<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploration of known and unknown early symptoms of cervical cancer and development of a symptom spectrum - Outline of a data and text mining based approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Ehrentraut</string-name>
          <email>ehrentraut@dsv.su.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karin Sundström</string-name>
          <email>karin.sundstrom@ki.se</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hercules Dalianis</string-name>
          <email>hercules@dsv.su.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer and Systems Sciences, (DSV) Stockholm University</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Medical Epidemiology and Biostatistics, (MEB) Karolinska Institutet</institution>
          ,
          <addr-line>Stockholm</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>34</fpage>
      <lpage>44</lpage>
      <abstract>
        <p>This position paper delineates the structure of some experiments to detect early symptoms of cervical cancer. We are using a large corpora of electronic patient records texts in Swedish from Karolinska University Hospital from the years 2009-2010, where we extracted in total 1,660 patient records with the ICD-10 diagnosis code C53 for cervical cancer. We used a Named Entity Recogniser called Clinical Entity Finder to detect the diagnosis and symptoms expressed in these clinical texts containing in total 2,988,118 words. We found 28,218 symptoms and diagnoses on these 1,660 patients. We present some initial findings, and discuss them and propose a set of experiments to find possible early symptoms and/or a spectrum of early symptoms of cervical cancer.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction and Motivation</title>
      <p>In the last ten years patient records have become, at least in Sweden, completely
digitalized and also centralised in large repositories. This is a vast source of
knowledge within medical research, however, this resource has not been much
exploited. The reason is that clinical researchers have little or no knowledge in
data and text mining, and also that these repositories due to their sensitive
nature are difficult to access in order to perform research.</p>
      <p>
        Lately, these sources have become to a very small extent available to
researchers in the U.S. as well as Europe. Meystre et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], wrote a review article
about different text mining approaches and tools, mostly for English textual
data. Among others, they mention an approach to detect early symptoms of
breast cancer. Dalianis et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] describes clinical text mining including
extraction and retrieval specifically for use in Swedish patient records. It is timely
to assess to what extent text mining can assist in the evaluation of symptoms
encountered in the course of human cancer.
      </p>
      <p>
        It has been shown in an interview-based study that young females with
cervical cancer frequently delay presentation, and not recognising symptoms as
serious may increase the risk of delay [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Improved identification and awareness
of early signs of cervical cancer may reduce both patient and provider delay of
investigation and treatment. Thus, it could be highly valuable to establish the
cervical cancer symptom spectrum and whether there are additional symptoms
that should be added to this. Text mining could as a novel tool aid with a
biasfree search of words and biochemical features that may not have been previously
suspected/identified by patients or health care.
      </p>
      <p>
        The hypothesis in this project originates from the assumption that women
with early cervical cancers and pre-cancers usually have no symptoms [
        <xref ref-type="bibr" rid="ref1 ref14">14, 1</xref>
        ].
So far, symptoms of a disease are mainly collected by means of capturing
descriptions made by the patient spontaneously, or after being questioned by a
health care professional. However, few of these are relevant for registration in
national health registers. Thus, traditional register-based research cannot access
such data.
      </p>
      <p>The project has two major aims:
1. Determine whether there are unknown early symptoms of cervical cancer,
and if so which. This to potentially inform health care and screening
processes of symptoms in women that may be of note. The anticipated output
is to find unknown early symptoms of cervical cancer. In this regard, a list
of concrete symptoms is considered to be the desired finding.
2. Develop and characterize a symptom spectrum for cervical cancer through a
holistic description of symptoms as recorded in medical text by health care
staff. Such a spectrum would include both previously known, and potentially
unknown symptoms. The anticipated output is a holistic description of
cervical cancer symptoms, i.e., likelihood of occurrence, time of occurrence and
frequency of occurrence of diverse symptoms. Ideally, the symptom
description will be an interactive visualization, as for instance depicted in Figure
1. This serves the purpose of generating a better understanding of possible
cervical cancer symptoms due to their potentially ambiguous nature.</p>
      <p>The purpose of both aims is to obtain a more concise understanding of
symptoms that occur in cervical cancer patients compared to non-cancer patients,
based on evidence that is gained through a statistical analysis of a large amount
of medical data. We intend to approach these aims by applying and enhancing
state of the art text mining tools.</p>
      <p>The overall goal is to use our findings as a complement in screening programs
for cervical cancer. In addition to taking a screening test for cervical cancer, the
physician could for example be able to run a program to filter out the patient’s
symptoms, if captured in the medical record, and compare them to a list of
possible early cervical cancer symptoms. Ideally, this approach should be generic
in order to be applicable to other cancers.</p>
      <p>This paper intends to outline the current state of the art within cervical
cancer prevention and how text mining is hitherto applied in the cancer domain.
Further, this paper presents (1) initial experiments that have been performed
as well as (2) an outlook on proposed work in order to find unknown early
symptoms and develop a symptom spectrum for cervical cancer.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        Cervical cancer (ICD-10 diagnosis code: C.53) is one of the most common cancers
worldwide [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], frequently striking young women below age 40, if not screened [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
A long-term infection with the Human papillomavirus (HPV), which spreads via
sexual contact, is deemed a necessary but not sufficient factor in the development
of cervical cancer [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Today, women are offered screening every three to five years, with the Pap
test being most commonly used, in order to detect abnormal changes in the cells
in an early stage. Cancer in an intermediate or advanced stage is highly mortal.
Early diagnosis is therefore crucial in order to prevent treatable pre-cancer from
turning into invasive cancer [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Early detection is yet often hindered since not
all women wish to participate in cervical screening programs.
      </p>
      <p>
        Women who do not attend screening can be diagnosed via symptomatic
presentation. However, diagnoses of cervical cancer may be delayed because of the
failure to recognize symptoms as cancer-related. As Lim et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] found, some
reasons for the delay may be that the patients (1) do not recognize possible
cancer symptoms, especially vaginal discharge, and (2) do not re-attend promptly
after first presentation despite persisting symptoms. Delays in diagnosis do also
occur on behalf of the provider who may fail to recognize cervical cancer-related
symptoms.
      </p>
      <p>
        According to the state of the art assumption, women with early cervical
cancers and pre-cancers usually have no symptoms [
        <xref ref-type="bibr" rid="ref1 ref14">14, 1</xref>
        ]. Yet, it is possible that
there are blood value deviations or other unforeseen symptoms. In most cases,
the symptoms do not start until the cancer has reached a more advanced stage.
Usual gynecological symptoms at that point are (1) abnormal vaginal bleeding,
(2) unusual discharge from the vagina and (3) pain during intercourse [
        <xref ref-type="bibr" rid="ref1 ref7">7, 1</xref>
        ].
      </p>
      <p>
        Increasing the awareness of (early) cervical cancer symptoms among women
and health care staff might improve diagnostics and chance of survival [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Finding hitherto unknown early symptoms which may appear during a pre-cancerous
stage could further help to diagnose cervical cancer at a time when it is still
treatable.
      </p>
      <p>
        Spasic et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] reviewed different approaches for clinical text mining within
the cancer domain. Of all studies the authors refer to, only two have focused on
cervical cancer and HPV, respectively.
      </p>
      <p>
        The study focusing on cervical cancer aimed at finding a method for retrieving
oncology documents relevant to clinical decision within the particular domain
of cervix cancer. With a content-based text classification process and similarity
analysis at its core, their system obtains its highest accuracy at 92% [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        The study focusing on HPV aimed at discriminating high-risk HPV types,
i.e., those that are related with cervical cancer, from low-risk types, i.e., those
that are not related with it. Comparing three machine learning algorithms,
namely AdaCost, AdaBoost and Naïve Bayes, the authors showed that
AdaCost outperforms the other algorithms, yielding an accuracy of circa 93% and
F-score of about 87% [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Materials and Methods</title>
      <p>
        The researchers of the MINECAN1 project and this particular study have
access to the Stockholm Electronic Patient Record (SEPR) Corpus that comprises
patient records from 2006 to 2014 from Karolinska University Hospital in
Stockholm, Sweden, [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The corpus contains records from all units at Karolinska
University Hospital except for records from the psychiatric and venereal disease
unit. For the MINECAN project, a subcorpus2 is created from the SEPR Corpus.
      </p>
      <p>In order to approach the main goal of finding unknown early symptoms and
creating a symptom spectrum for cervical cancer, the initial work comprised the
construction of part of the subcorpus and initial experiments performed on that
corpus.</p>
      <p>The approach used for this project resembles a retrospective case-control
study. That means past medical records are used to identify exposure and
outcome factors, e.g., potential exposures/symptoms for the outcome cervical
cancer. The study comprises a group of interest (study group) and a comparison (or
control) group3.
3.1</p>
      <sec id="sec-3-1">
        <title>ICD-10 diagnosis codes</title>
        <p>The study group data consists of records that belong to patients diagnosed with
cervical cancer. These patients are identified as having cervical cancer if an
appropriate ICD-10 diagnosis code is found in their records. All cervical cancer
related ICD-10 codes were specified by the project’s medical expert. They are:
1 MINECAN - Data and text mining of cancer symptoms and comorbidities in
electronic patient records in the Nordic languages
2 This research has been approved by the Regional Ethical Review Board in Stockholm
(Etikprövningsnämnden i Stockholm), permission number 2014/1882-31/5
3 http://hsl.lib.umn.edu/biomed/help/understanding-research-study-designs
– C53.0 (Malignant neoplasm: Endocervix)
– C53.1 (Malignant neoplasm: Exocervix)
– C53.8 (Malignant neoplasm: Overlapping lesion of cervix uteri)
– C53.9 (Malignant neoplasm: Cervix uteri, unspecified)
– D06.0 (Carcinoma in situ: Endocervix)
– D06.1 (Carcinoma in situ: Exocervix)
– D06.7 (Carcinoma in situ: Other parts of cervix)
– D06.9 (Carcinoma in situ: Cervix, unspecified)
– N87.2 (Severe cervical dysplasia, not elsewhere classified).</p>
        <p>The SEPR Corpus is stored in a database. Ultimately, the subset that is
created from this corpus for the cervical cancer project will comprise records that
belong to the study as well as as control group. As part of the first experiments,
only data for the study group has been extracted. Defining and extracting data
for the control group will be done at a later point in time.</p>
        <p>For the study group, the following information is extracted from the database
using MySQL queries:
– Gender and age of patient
– Date of patients’ admission to and discharge from hospital
– Clinic(s) where patient is treated
– Daily note (free text) and corresponding date of entrance into hospital system
during the years 2009-2010</p>
        <p>Once the data is extracted, all information about the patients is saved into
a text file with one file per patient, containing patient number, age and gender
information as well as all the patients’ daily notes sorted by date. These files are
then used for further processing and analysis.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Statistics of study group</title>
        <p>Statistics for the study group were obtained according to the following
parameters: age, clinic, time of diagnosis, length of treatment.</p>
        <p>In total, 1,660 patients are contained in the study group. Of these patients
1,587 patients have obtained only one ICD-10 diagnosis code, i.e., a C53, D06 or
N87 code. 72 patients have had two diagnosis codes in their records, 42 patients
had C53 and D06 diagnosis codes in their records while for 29 patients, D06 and
N87 co-occurred in the records. No patients had a C53 and N87 co-occurring in
the record. For one patient, all three diagnosis codes occurred in the record. Of
the 1,587 patients who only had one diagnosis, 603 had a C53 diagnosis code,
955 a D06 code and 29 a N87 code.</p>
        <p>The following section describes an initial approach of generating a frequency
list of symptoms captured in records of patients that were assigned a C53 code4
on a small subset of the data.
4 Only using C53 codes and no D06 and N87 codes is motivated by the fact that we
want to start testing</p>
        <p>This method aims at identifying symptom words in patient records, extract
them from the records and sort them according to their frequencies. Ultimately,
this step will be done for the records of the study group and the control group,
resulting in two frequency lists, a cervical cancer list and a control list. The two
frequency lists will be compared to one another to see
– if and how the symptoms differ between cases and controls
– if well-known cervical cancer symptoms are identified most frequently in the
cervical cancer list or
– if there are other symptoms that occur more frequently
– whether our methods can accurately identify a priori known/suspected
associations, which should validate whether the methodology is appropriate
As part of these first experiments, an initial cervical cancer frequency list
was created in a two step process.</p>
        <p>– Identify all symptoms by using the tool Clinical Entity Finder (CEF)
– Extract, sort and count all found symptoms and save them into a frequency
list</p>
        <p>
          The Clinical Entity Finder, CEF, implements the idea/task of Named
Entity Recognition (NER), i.e., recognizing expressions denoting entities such as
diseases, drugs, or people’s names in free text documents [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. This task can be
performed automatically and over the past years multiple NER algorithms have
been implemented. NER modules for English are for instance available via the
Stanford CoreNLP5 package or Apache OpenNLP6. Skeppstedt et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] have
5 http://nlp.stanford.edu/software/corenlp.shtml. 2014-09-08.
        </p>
        <p>6 https://opennlp.apache.org/index.html. 2014-09-08.
implemented the Clinical Entity Finder that can automatically recognize
entities in narrative text of Swedish health records. The tool is based on CRF++,
an implementation of the conditional random fields algorithm, and is initially
implemented to detect the terms within the entity categories Disorder, Finding,
Pharmaceutical Drug and Body Structure.</p>
        <p>After running CEF, the detected cervical cancer symptoms are sorted, counted
and saved into a frequency list that is depicted in Table 1.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>– Multiple inflectional forms of the same word, such as smärta (Engl.: pain)
and smärtor (Engl.: pains), occur in the frequency list. Using lemmatization,
they should be reduced to their base form in order to only include the main
symptom concept in the frequency list.
– So far the frequency list contains symptoms which are negated and that
should be removed from the list. Negation detection will need to be applied
in order to filter out these symptoms.
– Since we are interested in early symptoms, mainly daily patient notes that
are added to the EHR before the cancer diagnosis are of interest. So far we
used all patient notes that exist in the EHR for a patient with a cervical
cancer diagnosis. A future task aims at using only those notes made before
diagnosis, when detecting symptoms and generating frequency lists from
them.
– Identifying symptoms by applying CEF yielded promising results. Yet CEF
should be adapted to the domain by using more domain relevant training
data and incorporating negation.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>During our research work we encountered some challenges. We are not yet at the
stage of identifying any early unknown symptoms of cervical cancer but are able
to succesfully confirm other known symptoms such as bleeding that is a possible
symptom of cervical cancer.</p>
      <p>Some of the symptoms we identified were actually negated symptoms as not
bleeding, findings that our system could not identify as negated findings/symptoms,
since we did not use any negation detection system. Some of the symptoms which
are enumerated in Table 1. are therefore negated.</p>
      <p>Findings that are in singular or plural form as bleeding or bleedings could be
reduced to one base form using a lemmatizer. The same approach can be carried
out for determined and non-determined form of nouns. Determined nouns in
Swedish uses a inflection en to change to determined form; blödning+en =&gt;
blödning en, instead of a modifier as in English, the bleeding. Reducing these
identical findings would make the analyse easier and increase precision.</p>
      <p>Another obstacle was temporality, the patient record stretches over several
months or years and we need a method to identify when something occurred.
Certainly we have time stamps on each note, but within each note the physician
sometimes refer to earlier findings and relate to them.</p>
      <p>Regarding identifying terms we saw that there are many non-standard words
and abbreviations and compounds of abbreviations and words, as for example,
cervixca., that CEF could could not identify as named entities. This could easiest
be solved by adding in-domain annotated data.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and Further Work</title>
      <p>This paper described the first steps towards finding unknown early symptoms
and building a symptom spectrum for cervical cancer. As the projects progresses
we plan to work on the following tasks:</p>
      <p>– Defining and extracting the control group
– Testing and advancing the following methods to identify symptoms captured
in the patient records:</p>
      <sec id="sec-6-1">
        <title>NER and frequency counting approach</title>
        <p>Named Entity Recognition and Random Indexing</p>
        <p>Clustering
– Using and adapting existing text mining tools for the domain and
incorporating them into the preceding methods:</p>
      </sec>
      <sec id="sec-6-2">
        <title>Lemmatization</title>
        <p>Negation and certainty detection
Temporality</p>
        <p>Mapping symptoms to ICD-10 codes
– Analyzing and assembling the results as well as designing a visual
representation for the developed symptoms spectrum.</p>
        <p>One limitation may be that aim 1, finding previously early and/or unknown
symptoms of cervical cancer, cannot be fulfilled. However, this in turn could
actually inform health care practice and confirm the current evidence base for
cervical cancer as a relatively symptom-less disease, demonstrated by
systematically exploiting a novel data source; medical records. Regardless of aim 1, our
aim 2 should provide valuable information on the symptom spectrum in cervical
cancer.</p>
        <p>This paper has outlined the current state-of-the-art within cervical cancer
prevention and how text mining is hitherto applied in the cancer domain.
Further, this paper presented (1) initial experiments that have been performed as
well as (2) an outlook on proposed work in order to find unknown early symptoms
and develop a symptom spectrum for cervical cancer.</p>
        <p>We believe that outlining the scope of the project, including aims,
state-ofthe-art research, proposed future work and limitations, as well as performing
initial experiments was crucial for enabling a stringent work flow in the project.</p>
        <p>
          Our methodology can also been seen as a part of the HEALTH BANK
workbench proposed in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], that will offer processed aggregated and unaggregated
clinical data for research to be used in a wider context.
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>The authors would like to thank the Nordic Information for Action eScience
Center of Excellence in Health-Related e-Sciences (NIASC) and Nordforsk for
funding of the project and to Eric Thuning and Per "Pelle" Olofsson; both IT
experts at DSV for help with the management of the Stockholm EPR Corpus
server. We would also like to thank Maria Skeppstedt for letting us use the
Clinical Entity Finder and for Aron Henriksson for assisting us in executing
Clinical Entity Finder.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>American</given-names>
            <surname>Cancer Society</surname>
          </string-name>
          , A.:
          <article-title>Cervical Cancer Prevention</article-title>
          and
          <string-name>
            <given-names>Early</given-names>
            <surname>Detection</surname>
          </string-name>
          (
          <year>2014</year>
          ), http://www.cancer.org/acs/groups/cid/documents/webcontent/003094- pdf.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Cancer</given-names>
            <surname>Research</surname>
          </string-name>
          <string-name>
            <surname>UK</surname>
          </string-name>
          , U.:
          <article-title>Worldwide cancer incidence statistics</article-title>
          , http://www.cancerresearchuk.org/cancer-info/cancerstats/world/ incidence/Common, visited: November 13th 2014
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dalianis</surname>
          </string-name>
          , H.:
          <article-title>Clinical text retrieval - an overview of basic building blocks and applications</article-title>
          . In: Paltoglou,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Loizides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Hansen</surname>
          </string-name>
          , P. (eds.) Professional Search in the Modern World, vol.
          <volume>8830</volume>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>165</lpage>
          . Springer Verlag, Lecture Notes in Computer Science (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dalianis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hassel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henriksson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skeppstedt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Stockholm EPR Corpus:
          <article-title>A clinical database used to improve health care</article-title>
          .
          <source>In: Swedish Language Technology Conference</source>
          . pp.
          <fpage>17</fpage>
          -
          <lpage>18</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dalianis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henriksson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kvist</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weegar</surname>
          </string-name>
          , R.:
          <article-title>HEALTH BANK - A Workbench for Data Science Applications in Healthcare</article-title>
          .
          <source>In: Proceedings of CAiSE'15 - Industry Track. Springer Verlag, Lecture Notes in Computer Science</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Swedish</surname>
          </string-name>
          <article-title>Council on Health Technology Assessment</article-title>
          , SBU, S.:
          <article-title>Tidig upptäckt av symtomgivande cancer - En systematisk litteraturöversikt, (In Swedish)</article-title>
          ,
          <source>(January</source>
          <year>2014</year>
          ), http://www.sbu.se/upload/Publikationer/Content0/1/ Tidig_upptackt_symtomgivande_cancer_smf.pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramirez</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sasieni</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patnick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forbes</surname>
            ,
            <given-names>L.J.:</given-names>
          </string-name>
          <article-title>Delays in diagnosis of young females with symptomatic cervical cancer in england: an interview-based study</article-title>
          .
          <source>British Journal of General</source>
          Practice pp.
          <fpage>e602</fpage>
          -
          <lpage>e610</lpage>
          (
          <year>October 2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forbes</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raju</surname>
            ,
            <given-names>K.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramirez</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          :
          <article-title>Measuring the nature and duration of symptoms of cervical cancer in young women: developing an interview-based approach</article-title>
          .
          <source>BMC women's health 13(1)</source>
          ,
          <volume>45</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Meystre</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kipper-Schuler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hurdle</surname>
          </string-name>
          , J.:
          <article-title>Extracting information from textual documents in the electronic health record: a review of recent research</article-title>
          .
          <source>IMIA Yearbook of Medical Informatics</source>
          <volume>47</volume>
          ,
          <fpage>128</fpage>
          -
          <lpage>144</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>S.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hwang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Zhang, B.T.:
          <article-title>Mining the risk types of human papillomavirus (HPV) by AdaCost</article-title>
          . In: Mařík,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Retschitzegger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Štěpánková</surname>
          </string-name>
          ,
          <string-name>
            <surname>O</surname>
          </string-name>
          . (eds.)
          <source>Database and Expert Systems Applications</source>
          . Springer (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Polpinij</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Ontology-based text analysis approach to retrieve oncology documents from PubMed relevant to cervical cancer in clinical trials</article-title>
          .
          <source>In: ICDM Workshop on Advances in Data Mining. IBaI Publishing</source>
          , Leipzig (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Skeppstedt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kvist</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nilsson</surname>
            ,
            <given-names>G.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalianis</surname>
          </string-name>
          , H.:
          <article-title>Automatic recognition of disorders, findings, pharmaceuticals and body structures from clinical text: An annotation and machine learning study</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>49</volume>
          ,
          <fpage>148</fpage>
          -
          <lpage>158</lpage>
          (
          <year>June 2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Spasić</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Livsey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keane</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nenadić</surname>
          </string-name>
          , G.:
          <article-title>Text mining of cancer-related information: Review of current status and future directions</article-title>
          .
          <source>International journal of medical informatics 83(9)</source>
          ,
          <fpage>605</fpage>
          -
          <lpage>623</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Storck</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Cervical dysplasia</article-title>
          .
          <source>Online</source>
          (
          <year>2014</year>
          ), http://www.nlm.nih.gov/ medlineplus/ency/article/001491.htm, medlinePlus
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Sundström</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Human Papillomavirus Test and Vaccination - Impact on Cervical Cancer Screening and Prevention</article-title>
          .
          <source>Ph.D. thesis</source>
          , Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, Sweden (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Walboomers</surname>
            ,
            <given-names>J.M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacobs</surname>
            ,
            <given-names>M.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manos</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosch</surname>
            ,
            <given-names>F.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kummer</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>K.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snijders</surname>
            ,
            <given-names>P.J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peto</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meijer</surname>
            ,
            <given-names>C.J.L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muñoz</surname>
          </string-name>
          , N.:
          <article-title>Human papillomavirus is a necessary cause of invasive cervical cancer worldwide</article-title>
          .
          <source>The Journal of Pathology</source>
          <volume>189</volume>
          (
          <issue>1</issue>
          ),
          <fpage>12</fpage>
          -
          <lpage>19</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>