<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Task 1a of the CLEF eHealth Evaluation Lab 2015</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Clinical Speech Recognition</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hanna Suominen</string-name>
          <email>hanna.suominen@nicta.com.au</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leif Hanlen</string-name>
          <email>leif.hanlen@nicta.com.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <email>lorraine.goeuriot@imag.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <email>liadh.kelly@scss.tcd.ie</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J F Jones</string-name>
          <email>Gareth.Jones@computing.dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dublin City University</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NICTA</institution>
          ,
          <addr-line>ANU, and UC, Canberra, ACT</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>NICTA, The Australian National University (ANU), University of Canberra (UC), and University of Turku</institution>
          ,
          <addr-line>Canberra, ACT</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Trinity College Dublin</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Université Grenoble Alpes</institution>
          ,
          <addr-line>Grenoble</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Best practice for clinical handover and its documentation recommends standardized, structured, and synchronous processes with patient involvement. Cascaded speech recognition (SR) and information extraction could support their compliance and release clinicians' time from writing documents to patient interaction and education. However, high requirements for processing correctness evoke methodological challenges. First, multiple people speak clinical jargon in the presence of background noise with limited possibilities for SR personalization. Second, errors multiply in cascading and hence, SR correctness needs to be carefully evaluated as meeting the requirements. This overview paper reports on how these issues were addressed in a shared task of the eHealth evaluation lab of the Conference and Labs of the Evaluation Forum in 2015. The task released 100 synthetic handover documents for training and another 100 documents for testing in both verbal and written formats. It attracted 48 team registrations, 21 email confirmations, and four method submissions by two teams. The submissions were compared against a leading commercial SR engine and simple majority baseline. Although this engine performed significantly better than any submission [i.e., 38.5 vs. 52.8 test error percentage of the best submission with the Wilcoxon signed-rank test value of 302.5 (p &lt; 10-12)], the releases of data, tools, and evaluations contribute to the body of knowledge on the task difficulty and method suitability. Contributor Statement: HS, LH, and GJFJ designed the task and its evaluation methodology. HS developed the dataset and together with LH, led the task as a part of the CLEFeHealth2015 evaluation lab, chaired by LG and LK. HS drafted the paper and after this all authors revised and approved it.</p>
      </abstract>
      <kwd-group>
        <kwd>Computer Systems Evaluation</kwd>
        <kwd>Data Collection</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Medical Informatics</kwd>
        <kwd>Nursing Records</kwd>
        <kwd>Patient Handoff</kwd>
        <kwd>Patient Handover</kwd>
        <kwd>Records as Topic</kwd>
        <kwd>Software Design</kwd>
        <kwd>Speech Recognition</kwd>
        <kwd>Testset Generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Fluent information flow, defined as channels, contact, communication, or links to
pertinent people [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], is critical in healthcare in general and in particular in clinical
handover (aka handoff), when a clinician or group of clinicians is transferring
professional responsibility and accountability, for example, at shift change of nurses [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
This shift-change nursing handover is a form of clinical narrative where only a small
part of the flow is documented in writing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Best practice recommends standardized,
structured, and synchronous processes for handover and its information
documentation not only in the presence but also in active involvement of the patients, and where
relevant, their next-of-kin [
        <xref ref-type="bibr" rid="ref5">4</xref>
        ].1 However, failures in information flow from nursing
handover are a major contributing factor in over two-thirds of sentinel events in
hospitals and associated with over a tenth of preventable adverse events [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Only after a
couple of shift changes, anything from two-thirds to all verbal handover information
is lost or, even worse, transferred incorrectly if not documented electronically in
writing [
        <xref ref-type="bibr" rid="ref6 ref7">5, 6</xref>
        ].
      </p>
      <p>
        In order to support compliance with these processes, cascaded speech recognition
(SR) with information extraction (IE) has been studied in 2015 [
        <xref ref-type="bibr" rid="ref8 ref9">7, 8</xref>
        ]. As justified
empirically in clinical settings in 2014, the cascade pre-fills a structured handover
form for a clinician to proof and sign off [
        <xref ref-type="bibr" rid="ref10 ref11">9, 10</xref>
        ]. Based on the aforementioned rate of
information loss, the approach of the nurse who is handing over proofing and signing
off the document draft him/herself any time before the shift ends (but preferably
immediately after the handover) can decrease the loss to 0–13 per cent.
      </p>
      <p>
        This novel application evokes fruitful challenges for method research and
development, and consequently, its first part (i.e., clinical SR) was chosen as the Task 1a of
the eHealth Evaluation Lab by the Conference and Labs of the Evaluation Forum
(CLEF) in 2015 [
        <xref ref-type="bibr" rid="ref12">11</xref>
        ].2 First, clinical characteristics complicate SR. This derives from
a large number of nursing staff moving between patient-sites to involve patients in
handover, resulting in a noisy minimally-personalized multi-speaker setting far from a
typical case with a single person, equipped with a personalized SR engine, speaking
in a peaceful office. Second, SR errors multiply in cascading and, because of the
severe implications that they may have in clinical decision-making, the cascade
correctness needs to be carefully evaluated as meeting the clinical requirements.
      </p>
      <p>
        The task aligns with the CLEFeHealth usage scenario of easing patients, their
nextof-kin, and other laypersons in understanding and accessing electronic health
(eHealth) information [
        <xref ref-type="bibr" rid="ref13 ref14">12, 13</xref>
        ]. Namely, the application could release a substantial
amount of nurses’ time from documentation to, for example, longer discussions about
the findings, care plans, and consumer-friendly resources for further information with
1 Also the World Health Organisation (WHO) provides similar guidance as a mechanism to
contribute to safety and quality in healthcare at
http://www.who.int/patientsafety/research/methods_measures/hu
man_factors/organizational_tools/en/ (all websites of this paper were
accessible on 25 May 2015)
2 https://sites.google.com/site/clefehealth2015/
the patients, and where relevant, their next-of-kin. Documenting every event in
healthcare, as required by law, can take nearly sixty per cent of nurses’ working time
with centralized clinical information systems or fully structured information entry
(whilst free-form text entry at the patient-site decreases this to a few minutes per
patient) [
        <xref ref-type="bibr" rid="ref15 ref16 ref17">14–16</xref>
        ]. For example, every year within the Organisation for Economic
Cooperation and Development (OECD), on average seven physician consultations and
0.2 hospital discharges take place per capita.3 SR writes a document draft from a tenth
to three-quarters of the time it takes to transcribe this by hand, whilst the clinician’s
proofing time is about the same in both cases [
        <xref ref-type="bibr" rid="ref18">17</xref>
        ]. The speech-recognized draft for a
minute of verbal handover (with 160 words, corresponding to the range that people
comfortably hear and vocalize words [
        <xref ref-type="bibr" rid="ref19">18</xref>
        ]) is available only 20 seconds after finishing
the handover with a real-time engine that recognizes at least as many words per
minute as a very skilled typist (i.e., 120 [19]). Cascading this with content structuring
through IE can bring further efficiency gains by easing finding information and
making this content available for computerized decision-making and surveillance in
healthcare [20].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Materials and Methods</title>
      <p>The hold-out method was used in performance evaluation of the task. Task materials
consisted of a training set of 100 synthetic patient cases and an independent set of
another 100 synthetic patient cases for testing. Given the training set, the task was to
minimize the number of incorrectly recognized words on the test set (i.e., on the
heldout set). Performance of the submitted methods and two baseline methods was
compared statistically.
2.1</p>
      <sec id="sec-2-1">
        <title>Dataset for Training</title>
        <p>
          The dataset called NICTA Synthetic Nursing Handover Data was used in this task for
method development and training [
          <xref ref-type="bibr" rid="ref9">8</xref>
          ].4 This set of 100 synthetic patient cases was
developed for SR and IE related to nursing shift-change handover in 2012–2014. Each
case consisted of a patient profile; a written, free-form text paragraph (i.e., the written
handover document) to be used as a reference standard in SR; and its spoken (i.e., the
verbal handover document) and speech-recognized (i.e., speech-recognized
documents with respect to six vocabularies) counterparts. The dataset was released on the
task page on 15 November 2014.
        </p>
        <p>First, the first author of this paper (Adj/Prof in machine learning and
communication for health computing) generated 100 synthetic patient profiles, using common
user profile generation techniques [21]. With an aim for balance in patient types, she
created profiles for 25 cardiovascular, 25 neurological, 25 renal, and 25 respiratory
3 Derived from OECD.StatsExtracts (http://stats.oecd.org/) for 2009 (i.e., the
most recent year that has almost all data available)
4
http://www.nicta.com.au/nicta-synthetic-nursing-handoveropen-data-software-and-demonstrations/
patients of an imaginary medical ward for adults in Australia. These patient types
were chosen because they represent the most common chronic diseases and national
priority areas [22]. The reason for patient admission was always an acute condition,
but some patients had also chronic diseases. Some patients were recently admitted to
the ward, some had been there for some days already, and some were almost ready to
be discharged after a shorter or longer inpatient period. Each profile was saved as a
DOCX file and contained a stock photo from a royalty-free gallery, name, age,
admission story, in-patient time, and familiarity to the handover nurses.</p>
        <p>Second, the first author supervised a registered nurse (RN) in creating the written
handover documents for these 100 profiles. The RN had over twelve years’
experience in clinical nursing. Australian English was her second language and she was
originally from the Philippines. She was guided to imagine herself working in the
medical ward and delivering verbal shift-change handovers to another nurse by the
patient’s bedside as if she was talking. All handover information was to be given as a
100–300-word monologue, using normal wordings. The resulting realistic but fully
imaginary handovers were saved as TXT files.</p>
        <p>
          Third, the first author supervised the RN in creating the verbal handover
documents by reading the written handover documents out loud as the nurse giving the
handover. She was guided to record in a quiet office environment, try to speak as
naturally as possible, avoid sounding like reading text, and repeat the take until she
was satisfied with the outcome. The Olympus WS-760M digital recorder [purchased
for 269.00 AUD (191 €) in October 2011, weight of 51 g, dimensions of 98.5 mm ×
40.0 mm × 11.0 mm] and Olympus ME52W noise-canceling lapel-microphone
[purchased for 15.79 AUD (11 €) in October 2011, weight of 4 g (+ 11 g for an optional
cable and clip), dimensions of 35 mm (+ a 15 mm plug) × 13 mm × 13 mm] were
used, because they were previously shown to produce superior word correctness in SR
[
          <xref ref-type="bibr" rid="ref8">7</xref>
          ].5, 6 Each document was saved as a WMA file and then converted from stereo to
mono tracks and exported as WAV files on Audacity 2.0.3 for Mac.7
        </p>
        <p>
          Fourth, the first author used Dragon Medical 11.0 for SR.8 This clinical engine was
chosen because it included an option for Australian English. It was first initialized
with not only this accent but also to the RN’s age of 22–54 years and recording of The
Final Odyssey [DOCX file of 3,893 words in writing and WMA file of 29 minutes 22
seconds as speech (4 minutes needed)] using the aforementioned recorder and
microphone. Also these training/personalization/adaption files were released. Six Dragon
vocabularies (i.e., general as the most generic clinical vocabulary, medical because of
the medical ward, nursing because of the nursing handover, cardiology because of the
cardiac patients, neurology because of the neurological patients, and pulmonary
disease because of the respiratory patients) were compared although the nursing
vocabu5 http://www.olympus.co.uk/site/en/archived_products/audio/
audio_recording_1/ws_760m/index.pdf
6
https://shop.olympus.eu/UK-en/microphones/olympus/me52w-minimono-microphone-p-239.htm
7 http://sourceforge.net/projects/audacity/
8
http://www.nuance.com/products/dragon-medical-practiceedition/index.htm
lary shown to produce the best results in SR [
          <xref ref-type="bibr" rid="ref8 ref9">7, 8</xref>
          ]. Each speech-recognized document
was saved as a TXT file.
        </p>
        <p>
          The data release with the requirement to cite [
          <xref ref-type="bibr" rid="ref9">8</xref>
          ] was approved at NICTA and the
RN was consented in writing. The license of the verbal, free-form text documents
(i.e., WMA and WAV files) was Creative Commons - Attribution Alone -
Noncommercial - No Derivative Works for the purposes of testing SR and language
processing algorithms.9 The remaining documents (i.e., DOCX and TXT files) were
licensed under Creative Commons – Attribution Alone.10
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>An Independent Dataset for Testing</title>
        <p>The training set was supplemented with an independent dataset for testing. This
additional set of 100 synthetic patient cases was developed in 2015. Each case consists of
(1) a patient profile, (2) a written handover document, (3) a verbal handover
document, and (4) a speech-recognized document with respect to the nursing vocabulary.
Its subset of documents (3) and (4) was released on the task page on 23 April 2015;
the organizers did not release the profiles in order to avoid their contents to be used an
a processing input, and they held the written handover documents out as a blind set
for anyone but the first author and RN to ensure independent training and testing. The
set was created the same way as the training set except that the profile photos were
reused, software was updated to Audacity 2.1.3 for Mac with ffmpeg-mac-2.2.2,11 and
only the nursing vocabulary was chosen for SR.</p>
        <p>The data release was approved at NICTA and the RN was consented in writing.
The licensing constraints are the same as before; however, we ask to cite this task
overview for the data release.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Submission to Performance Evaluation</title>
        <p>The participants needed to submit their processing results by 1 May 2015 using the
Easy Chair System of the lab.12 Submissions that developed the SR engine itself were
evaluated separately from those that studied post-processing methods for the
speechrecognized text. Also a separate submission category was assigned to solutions based
on both SR and text post-processing.</p>
        <p>Only fully automated methods were allowed, that is, human-in-the-loop means
were not permitted. Each participant was allowed to submit up to two
methods/parameterizations/compilations (referred to as a method from now on) to the first
category and up to two methods to the second category. If addressing both these
categories, the participant was asked to submit all possible combinations of these methods
as their third category submission (i.e., up to 2 × 2 = 4 files).
9 http://creativecommons.org/licenses/by-nc-nd/4.0/
10 http://creativecommons.org/licenses/by/4.0/
11 https://www.ffmpeg.org/
12 https://easychair.org/conferences/?conf=clefehealth2015resul
with the professional license</p>
        <p>In addition to the submission category, the submissions consisted of the following
elements: team name and description; address of correspondence; author(s); at least
three method keywords;13 method description (max 100 words per method); and
processing outputs for each method on the 100 training and 100 test documents.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Methods and Measures in Performance Evaluation</title>
        <p>We challenged the participants to minimize the number of incorrectly recognized
words on the independent test set. This correctness was evaluated on the entire test set
using the primary measure of the percentage of incorrect words [aka the error rate
percentage (E)] as defined by the Speech Recognition Scoring Toolkit (SCTK), 2.4.0
without punctuation as a differentiating feature.14 This measure sums up the
percentages of substituted (S), deleted (D), and inserted (I) words (i.e., E = S + D + I) and
consequently, the smaller the value of E, the better the performance. To illustrate
these error types, speech-recognizing your word as you are had the substitution (your,
you), insertion are, and deletion word.</p>
        <p>As secondary measures, we reported the percentage of correctly detected words
(C) on the entire test set together with the breakdown of E to D, I, and S. We also
documented the raw numbers of correct (nC), substituted (nS), deleted (nD), and
inserted words (nI). Notice that C + S + D = 100 and nC + nS + nD is the number of words in
the reference standard.</p>
        <p>To provide more details on performance differences across the individual handover
documents, we also computed the error rate percentage ei in individual documents di
∈ {d1, d2, d3, ... , d100} of the test set. Then, we summarized these values through their
minimum (min), maximum (max), mean, median, and standard deviation (SD).</p>
        <p>Finally, instead of evaluating this generalization capability of the method to unseen
data, we assessed the resubstitution performance on the training set; a method that
does not perform well even on its training set is poor, but excellence on training set
may indicate over-fit, leading to issues in the generalizability.
2.5</p>
      </sec>
      <sec id="sec-2-5">
        <title>Baselines Methods in Performance Evaluation</title>
        <p>We used two baseline systems in the task, namely Dragon and Majority. The Dragon
baseline was based on Dragon Medical 11.0 with the nursing vocabulary and
initialization to the RN, recorder, and microphone. This commercial system included
substantial but closed domain dictionaries, had a substantial license fee per person
[purchased for 1,600.82 AUD (1,139 €) in January 2013], and was limited to the
Microsoft Windows operating system. Notice that the 200 synthetic handover cases were
not used to train the Dragon baseline. The Majority baseline assumed that the right
number of words is detected (i.e., the number of test words originating from the
refer13 preferably Medical Subject Headings
(http://www.nlm.nih.gov/mesh/MBrowser.html) or Association for
Computing Machinery classes (http://www.acm.org/about/class/ccs98-html)
14 http://www.itl.nist.gov/iad/mig/tools/
ence standard) and recognized each word as the most common training word (i.e.,
and) with correct capitalization (i.e., and for and, And for And, and so forth).</p>
        <p>To supplement the usage guidelines of SCTK,15 we provided the participants some
helpful tips (Appendix 1): we released an example script for removing punctuation
and formatting text files; a formatted reference file and Dragon baseline for the
training set; overall and document-specific evaluation results for this file pair; and
commands to perform these evaluations and ensure the correct installation of SCTK.
2.6</p>
      </sec>
      <sec id="sec-2-6">
        <title>Statistical Significance Testing</title>
        <p>
          Statistical differences between the error rate percentages of the two baselines and
participant submissions were evaluated using the Wilcoxon signed-rank test (W) [
          <xref ref-type="bibr" rid="ref4">23</xref>
          ].
This test was chosen as an alternative to the paired t-test, because the Shapiro-Wilk
test [24, 25] with the significance level of 0.005 indicated that the error rate
percentages ei for the sample of the 100 test documents were not normally distributed (e.g., p
values of 0.018 and 0.225 for the Dragon and Majority baselines, respectively).
        </p>
        <p>After ranking the baselines and submissions based on their error rate percentage on
the entire dataset for testing, W was computed for the paired comparisons from the
best and second-best method to the second-worst and worst method. The resulting p
value and the significance level of 0.05 was used to determine if the median
performance of the higher-ranked method was significantly better than this value for the
lower-ranked method. All statistical tests were computed using R 3.2.0.16
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>The task released in both verbal and written formats the total of 200 synthetic clinical
documents that can be used for studies on nursing documentation and informatics. It
attracted nearly 50 team registrations with about half of them confirming their
participation through email. Two teams submitted two SR methods each. Although no
method performed as well as the Dragon baseline, the task contributed to the body of
knowledge on the task difficulty and method suitability.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Release</title>
        <p>The task released a training set of 100 documents (Fig. 1) on 27 October 2014; an
independent test set of 100 documents on 23 April 2015; and reference standard for
the test documents together with processing results of the Dragon baseline and
subKen harris, bed three, 71 yrs old under Dr Gregor, came in with arrhythmia. He complained
of chest pain this am and ECG was done and was reviewed by the team. He was given some
15
http://www1.icsi.berkeley.edu/Speech/docs/sctk</p>
        <p>1.2/options.htm#option_r_name_0
16 http://www.r-project.org/
anginine and morphine for the pain. Still tachycardic and new meds have been ordered in the
medchart. still for pulse checks for one full minute. Still awaiting echo this afternoon. His BP
is just normal though he is scoring MEWS of 3 for the tachycardia. He is still for monitoring.
Dragon baseline: Own now on bed 3 he is then Harry 70 is 71 years old under Dr Greco he
came in with arrhythmia he complained of chest pain this morning in ECG was done and
reviewed by the team he was given some and leaning in morphine for the pain in she is still
tachycardic in new meds have been ordered in the bedtime is still 4 hours checks for one full
minute are still waiting for echocardiogram this afternoon he is BP is just normal though he
is scarring meals of 3 for the tachycardia larger otherwise he still for more new taurine
missions later in 2015.17 Errors that the Dragon baseline made on training documents
ten or more times included substituting years for yrs (n = 48), in with and (n = 22),
one with 1 (n = 17), alos with obs (n = 12), and to with 2 (n = 12); deleting is (20),
are (13), and and (11); and inserting and (210), is (136), in (106), she (71), are (58),
all (45), arm (44), for (43), the (37), he (35), that (34), a (27), her (19), eats (15), on
(15), also (14), am (12), does (11), bed (10), s (10), and to (10) [26].</p>
        <p>The training set had 7,277 words and 1,304 (1,377) of them were unique without
(with) capitalization as a differentiating feature. For the test set, these numbers were
6,818, 1,279, and 1,323, respectively. Although both sets followed the Zipf’s law [27]
(Fig. 2), and were thereby typical language samples, their vocabularies shared only
645 (738) unique words without (with) capitalization as a differentiating feature (Fig.
3). Consequently, the sets can be seen as fairly independent, as intended. The 10
common words without capitalization as a differentiating feature were and (347
occurrences in the training asset and 315 in the test set), is (256, 288), he (243, 201), in
(170, 212), for (163, 140), with (162, 151), she (151, 152), on (141, 175), the (138,
88), and to (124, 73).
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Community Interest and Participation</title>
        <p>The task was open for everybody. We particularly welcomed academic and industrial
researchers, scientists, engineers and graduate students in SR, natural language
processing, and biomedical/health informatics. We also encouraged participation by
multi-disciplinary teams that combine technological skills with nursing expertise.</p>
        <p>By 30 April 2015, 48 people had registered their interest in the task through the
CLEF 2015 registration system,18 and 21 of these team leaders had emailed to
confirm their participation. From its opening on 27 October 2014 to this date, the task
discussion forum gained five other members than the organizers.19</p>
        <p>By 1 May 2015, two teams submitted four methods. The first team, called
TUC_MI/MC, was from the Technische Universität Chemnitz (TUC) in Germany. Its
members were two researchers from the field of computer science, supervised by two
TUC professors. They followed an interdisciplinary approach where one part brought
the expertise from the field of speech processing to develop strategies for web-based
language model adaptation. The other one came from the field of information retrieval
to choose and develop methods for selecting and processing web resources to build a
thematically coherent adaptation corpus. The second team, called UC, was from the
University of Canberra in the Australian Capital Territory. It consisted of two PhD
17
http://www.nicta.com.au/nicta-synthetic-nursing-handoveropen-data-software-and-demonstrations/.
18 http://clef2015-labs-registration.dei.unipd.it/
19
https://groups.google.com/forum/#!forum/clefehealth2015-task1a-speech-recognition
students and three Professors from multi-disciplinary backgrounds, including clinical,
public health, machine learning, and software engineering, working in collaboration.20</p>
        <p>TUC_MI/MC submitted two SR methods. Their approach assumed each document
having its own context and hence suggested adapting SR for each document
separately. They used a two-pass decoding strategy: First, a verbal document was speech
recognized. Then, keywords of the utterances were extracted and used as queries in order
to retrieve web resources as adaptation data to build a document-specific dictionary
and language model with the interpolation weights of 0.8 and 0.9 for TUC_MI/MC.1
and TUC_MI/MC.2, respectively. Finally, re-decoding of the same document was
performed using the adapted dictionary and language model.</p>
        <p>Also UC submitted two SR methods. UC.1 was based on acoustic modeling of
speech using Hidden Markov Models (HMM). The verbal documents were
preprocessed, including filtering, word level segmentation, and Mel Frequency Cepstral
feature extraction, and HMM models were built for the training data. As there were
no repetitions of data from different sessions, bagging and bootstrapping of training
data were used. UC.2 combined language and acoustic models using the CMU Sphinx
open source toolkit for SR. A custom dictionary and language model was developed
for the speaker of the training set, because none of existing dictionary and language
models was suitable for her accent. Unfortunately, the organizers had to reject this
second method as even after an update request and deadline extension to 10 May
2015, this submission failed to meet the evaluation criteria for the format
and completeness.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Performance Benchmarks</title>
        <p>The Dragon baseline clearly had the best performance (i.e., E = 38.5) on the
independent set for testing, followed by the TUC_MI/MC.2 (E = 52.8), TUC_MI/MC.1 (E
= 52.3), UC.1 (E = 93.1), and the Majority baseline (E = 95.4) (Table 1). The
resubstitution performance of the first three methods was approximately the same (i.e.,
from E = 55.0 ± 0.9), but last two methods had nearly 100 per cent error also on the
training set (Table 2).</p>
        <p>The performance of the Dragon baseline on the test set was significantly better
than that of the second-best method (i.e., TUC_MI/MC.2, W = 302.5, p &lt; 10–12).
However, this rank-2 method was not significantly better than the third-best method
(i.e., TUC_MI/MC.1), but this rank-3 method was significantly better than the
fourthbest method (i.e., UC.1, W = 0, p &lt; 10–15). Finally, the performance of the
lowestranked method (i.e., the Majority baseline) was significantly worse than that of this
rank-4 method (W = 1,791.5, p &lt; 0.05).
20 Including Prof LH, task co-leader, as an advisor who encouraged participation without any
engagement in the team’s method experimentation and development. As noted in Section
2.2, he did not develop the test set nor had an access to it before other participants.
We conclude the paper by comparing the results with prior work, validating the
released data, and discussing the significance of this study.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Comparison with Prior Work</title>
        <p>
          SR at its best can achieve an impressive C of 90–99 with only 30–60 minutes of
personalization or adaptation to a given clinician’s speech [
          <xref ref-type="bibr" rid="ref18">17</xref>
          ]. This SR correctness is
supported by studies on mainly North-American male physicians speaking medical
documents. For studio recordings of a Spanish-accented Australian female nurse,
native Australian female nursing professional, and native Australian male physician
speaking nursing handover simulations, (C, E) pairs of Dragon Medical 11.0 are (62,
40), (64, 39), and (71, 32), respectively [
          <xref ref-type="bibr" rid="ref8">7</xref>
          ]. These numbers are very similar to those
for the Dragon baseline on the training set [i.e., (72, 56)] and test set [i.e., (73, 39)].
Differences between commercial engines (i.e., IBM ViaVoice 98, General Medicine
with C = 92 ± 1, L&amp;H Voice Xpress for Medicine 1.2, General Medicine with C = 86
± 1, and Dragon Medical 3.0 with C from 85 to 86) are not drastic [28]. To compare
this automation with the upper baseline of human performance, each clinical
document has 0.4 errors on average if transcribing by hand whilst for a speech-recognized
document, this number is 6.7 [29].
        </p>
        <p>We have studied correcting SR errors through post-processing in [26]. This
approach is unsupervised and applies phonetic similarity to substituted words. Its
evaluation on the 100 training documents gives promising results; in 15 per cent of all
1,187 unique substitutions by the Dragon baseline, the speech-recognized word
sounds exactly the same as its reference word and 23 per cent of them are at least 75
per cent similar.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Data Validation</title>
        <p>
          A basic scientific principle of the reproducibility of the results relies on availability of
open data, open source code, and open evaluation results [30, 31]. Access to research
data also increases the returns from public investment in this area; encourages
diversity of studies and opinion; enables the exploration of new topics and areas; and
reinforces open scientific inquiry [32]. Whilst this open movement in health sciences and
informatics is progressing, particularly for source code [33] and evaluation results
from clinical trials [
          <xref ref-type="bibr" rid="ref20">34</xref>
          ], its slowness in releasing data has significantly hindered
method research, development, and adoption [
          <xref ref-type="bibr" rid="ref21">35</xref>
          ]. Evaluation labs have improved the
situation [
          <xref ref-type="bibr" rid="ref21 ref22">35, 36</xref>
          ], but with some exceptions [
          <xref ref-type="bibr" rid="ref23 ref24">37, 38</xref>
          ],21 most open data are
deidentified [
          <xref ref-type="bibr" rid="ref13 ref14">12, 13</xref>
          ] and/or with use restriction [
          <xref ref-type="bibr" rid="ref14 ref25">39, 13</xref>
          ].22
        </p>
        <p>
          However, data de-identification on text documents is fraught with difficulties [
          <xref ref-type="bibr" rid="ref26">40</xref>
          ]
and the resulting data may still have some identifiable components [
          <xref ref-type="bibr" rid="ref27">41</xref>
          ].
Consequently, the minimum standard of clinical data de-identification recommend against
releasing verbal clinical documents or their transcriptions [
          <xref ref-type="bibr" rid="ref28">42</xref>
          ] – that is, precisely our open
data. Furthermore, de-identification is to be avoided on Australian clinical data,
because under Australian privacy law, it actually results in re-identifiable data, which
must have restricted use, appropriate ethical use, and approval from all data subjects
(e.g., patients, their visitors, nurses and other clinicians in the case of Australian
nursing shift-change handover with nurses’ team meeting followed by a patient-site
meet21 Synthetic clinical documents have been used in the evaluation labs of the NII Test Collection
for Information Retrieval Systems (NTCIR) for Japanese medical documents in 2013
(http://mednlp.jp/medistj-en/) and 2014
(http://mednlp.jp/ntcir11/).
22 These clinical data originating from US healthcare services accessible to registered users on
a password-protected Internet site (i.e., PhysioNetWorks at
https://physionet.org/works/) after manual authorization and approval of the
data access and use policies (e.g., for lab-participation or scientific purposes only).
ing [43, p. 27]. These use restrictions are even more complicated for the Australian
handover, as real documents that are not re-identifiable, apply to our patient-site case,
and allow releasing and use without, for example, commercial restriction do
not exist.23
        </p>
        <p>
          Due to the lack of existing text corpora, that match the Australian clinical setting,
and due to the difficulty of providing ethically-sound open data, we have
compromised by providing synthetic data that closely matches the real data typically found in
a nursing shift-change. We have validated this matching by employing and
projectfunding clinical experts to confirm the typicality and compare the synthetic
documents with real data and related processing results [
          <xref ref-type="bibr" rid="ref10 ref11 ref18 ref8 ref9">9, 17, 10, 7, 8</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>Significance</title>
        <p>
          The significance of our open synthetic clinical data lies in supporting innovation and
decreasing barriers of method research and development. In particular for new
participants in SR, IE, and other automated generation or analysis of text documents, the
aforementioned barrier of data access costs money and time. The entry and
transaction are even more expensive, but given the required expertise in the field, only the
latter cost can be decreased substantially by simplified licensing through open data
movement [
          <xref ref-type="bibr" rid="ref30">44</xref>
          ]. This also resolves the barrier of data access.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This shared task was partially supported by the CLEF Initiative and NICTA, funded
by the Australian Government through the Department of Communications and the
Australian Research Council through the Information and Communications
Technology (ICT) Centre of Excellence Program. We express our gratitude to Maricel Angel,
Registered Nurse at NICTA, for helping us to create this dataset. Last but not least,
we gratefully acknowledge the participating teams’ hard work. We thank them for
their submissions and interest in the task.
19. Ayres, R. U.., Martinás, K.: 120 wpm for very skilled typist. In: On the Reappraisal of
Microeconomics: Economic Growth and Change in a Material World, p. 41. Edward Elgar
Publishing, Cheltenham, UK &amp; Northampton, MA, USA (2005)
20. Allan, J., Aslam, J., Belkin, N., Buckley, C., Callan, J., Croft, B., Dumais, S., Fuhr, N.,
Harman, D., Harper, D. J., Hiemstra, D., Hofmann, T., Hovy, E., Kraaij, W., Lafferty, J.,
Lavrenko, V., Lewis, D., Liddy, L., Manmatha, R., McCallum, A., Ponte, J., Prager, J., Radev, D.,
Resnik, P., Robertson, S., Rosenfeld, R., Roukos, S., Sanderson, M., Schwartz, R., Singhal,
A., Smeaton, A., Turtle, H., Voorhees, E., Weischedel, R., Xu, J., Zhai, C.: Challenges in
information retrieval., language modeling: Report of a workshop held at the Center for
Intelligent Information Retrieval, University of Massachusetts Amherst, September 2002. SIGIR
Forum 37(1), 31–47 (2003)
21. Kuniavsky, M.: Observing the User Experience: A Practitioner’s Guide to User Research.</p>
      <p>Morgan Kaufmann Publishers, San Fransisco, CA, USA (2003)
22. Australian Government, Department of Health. Chronic disease: Chronic diseases are leading
causes of death., disability in Australia,
http://www.health.gov.au/internet/main/publishing.nsf/Content/
chronic (last updated: 26 September 2012)
23. Wilcoxon, F.: Individual comparisons by ranking methods. Biometrics Bulletin 1(6), 80–83
(1945)
24. Shapiro, S. S., Wilk, M. B.: An analysis of variance test for normality (complete samples).</p>
      <p>Biometrika 52(3–4), 591–611 (1965)
25. Razali, N., Wah, Y. B.: Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov,
Lilliefors., Anderson-Darling tests. Journal of Statistical Modeling., Analytics 2(1), 21–33
(2011)
26. Suominen, H. and Ferraro, G.: Noise in speech-to-text voice: Analysis of errors and feasibility
of phonetic similarity for their correction. In Karimi, S. and Verspoor, K. (eds.) Proceedings
of the Australasian Language Technology Association Workshop 2013 (ALTA 2013), pp. 34–
42. Association for Computational Linguistics (ACL), Stroudsburg, Brisbane, QLD,
Australia. (2013)
27. Powers, D. M. W. Applications and explanations of Zipf’s law. In Powers, D. M. W. (ed.):
Proceedings of the Joint Conferences on New Methods in Language Processing and
Computational Natural Language Learning (NeMLaP3/CoNLL’98), pp. 151–160. ACL, Stroudsburg,
PA, USA (1998)
28. Devine, E. G., Gaehde, S. A., Curtis, A. C.: Comparative evaluation of three continuous
speech recognition software packages in the generation of medical reports. JAMIA 7(5), 462–
468 (2000)
29. Al-Aynati, M. M. and Chorneyko, K. A.: Comparison of voice-automated transcription and
human transcription in generating pathology reports. Achieves of Pathology and Laboratory
Medicine 127(6), 721–725 (2003)
30. Sonnenburg, S., Braun, M. L., Ong, C. S., Bengio, S., Bottou, L., Holmes, G., LeCunn, Y.,
Müller, K.-R., Pereira, F., Rasmussen, C. E., Rätsch, G., Schölkopf, B., Smola, A., Vincent,
P., Weston, J., Williamson, R. C.: The need for open source software in machine learning.</p>
      <p>Journal of Machine Learning 8, 2443–2466 (2007)
31. Pedersen, T.: Empiricism is not a matter of faith. Computational Linguistics 34(3), 465–470
(2008)
32. Organisation for Economic Development (OECD): OECD Principles and Guidelines for</p>
      <p>Access to Research Data from Public Funding. OECD, Danvers, MA, USA (2007)
33. Estrin, D. and Sim, I.: Open mHealth architecture: An engine for health care innovation.
Science 330(6005), 759–760 (2010)</p>
      <sec id="sec-4-1">
        <title>Running SCTK</title>
        <p>The command
bin/sclite -r reference.txt -h reference.txt trn -i spu_id -o all
should produce perfect results by using the formatted reference standard both as a reference
standard (-r reference.txt) and speech-recognized (or hypothesized -h) text. The
command
bin/sclite -r reference.txt -h dragon_nursing.txt -i spu_id -o all
uses the formatted reference standard and formatted Dragon output, and hence should produce
the aforementioned document-specific and overall evaluation results. The -i spu_id option
refers to the transcription (TRN, i.e., a TXT file where each paragraph captures a handover
document, followed by (documentID_V) with _V specifying the person whose voice/speech is
recognized) formatted input files (with the default extended American Standard Code for
Information Interchange encoding and the default GNU diff alignment) to pair the reference
standard with the speech-recognized documents. The -o all option results in not only the
evaluation results as percentages and raw numbers but also more details for analyzing correctly
and incorrectly detected text patterns.</p>
        <p>Removing punctuation and changing to ASCII</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Glaser</surname>
            ,
            <given-names>S. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zamanou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hacker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Measuring, interpreting organizational culture</article-title>
          .
          <source>Management Communication Quarterly (MCQ) 1</source>
          (
          <issue>2</issue>
          ),
          <fpage>173</fpage>
          -
          <lpage>198</lpage>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Tran</surname>
          </string-name>
          , D. T..,
          <string-name>
            <surname>Johnson</surname>
          </string-name>
          , M.:
          <article-title>Classifying nursing errors in clinical management within an Australian hospital</article-title>
          .
          <source>International Nursing Review</source>
          <volume>57</volume>
          (
          <issue>4</issue>
          ),
          <fpage>454</fpage>
          -
          <lpage>462</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Finlayson</surname>
            ,
            <given-names>S. G.</given-names>
          </string-name>
          , LePendu,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <surname>N. H.</surname>
          </string-name>
          :
          <article-title>Building the graph of medicine from millions of clinical narratives</article-title>
          .
          <source>Scientific Data</source>
          <volume>1</volume>
          ,
          <issue>140032</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>23</surname>
          </string-name>
          <article-title>The written documents captured in existing clinical systems in Australia are typically entered by a clerk later in the shift and hence are not nursing handover transcripts</article-title>
          .
          <source>In other works [7</source>
          ,
          <issue>9</issue>
          , 10],
          <article-title>we placed microphones on nurses to evaluate real (de-identified) data. Consequently, they cannot be released</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          4.
          <string-name>
            <surname>Australian</surname>
          </string-name>
          <article-title>Commission on Safety and Quality in Healthcare (ACSQHC): Standard 6: Clinical handover</article-title>
          .
          <source>In: National Safety and Quality Health Standards</source>
          , pp.
          <fpage>44</fpage>
          -
          <lpage>47</lpage>
          . ACSQHC, Sydney,
          <string-name>
            <given-names>NSW</given-names>
            ,
            <surname>Australia</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          5.
          <string-name>
            <surname>Pothier</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monteiro</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mooktiar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Shaw</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Pilot study to show the loss of important data in nursing handover</article-title>
          .
          <source>British Journal of Nursing</source>
          <volume>14</volume>
          (
          <issue>20</issue>
          ),
          <fpage>1090</fpage>
          -
          <lpage>1093</lpage>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          6.
          <string-name>
            <surname>Matic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salamonson</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Review: Bringing patient safety to the forefront through structured computerisation during clinical handover</article-title>
          .
          <source>Journal of Clinical Nursing</source>
          <volume>20</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>184</fpage>
          -
          <lpage>189</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          7.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Johnson,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Sanchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Sirel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Basilakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Hanlen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Estival</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Dawson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Capturing patient information at nursing shift changes: Methodological evaluation of speech recognition and information extraction</article-title>
          .
          <source>Journal of the American Medical Informatics Association (JAMIA) 22(e1)</source>
          ,
          <fpage>e48</fpage>
          -
          <lpage>e66</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          8.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          , Ferraro, G.:
          <article-title>Benchmarking clinical speech recognition and information extraction: New data, methods, and evaluations</article-title>
          .
          <source>JMIR Medical Informatics</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ),
          <year>e19</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          9.
          <string-name>
            <surname>Dawson</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basilakis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>A usability framework for speech recognition technologies in clinical handover: A preimplementation study</article-title>
          .
          <source>Journal of Medical Systems</source>
          <volume>38</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          10.
          <string-name>
            <surname>Johnson</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchez</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basilakis</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dawson</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen L</surname>
          </string-name>
          .
          <article-title>Comparing nursing handover and documentation: Forming one set of patient information</article-title>
          .
          <source>International Nursing Review 2014</source>
          <volume>61</volume>
          (
          <issue>1</issue>
          ),
          <fpage>73</fpage>
          -
          <lpage>81</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          11.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Névéol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2015</article-title>
          .
          <source>In: CLEF 2015 - 6th Conference and Labs of the Evaluation Forum, Lecture Notes in Computer Science (LNCS)</source>
          . Springer, Berlin Heidelberg, Germany (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          12.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salanterä</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elhadad</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pradhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>South</surname>
            ,
            <given-names>B. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mowery</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G. J. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the ShARe/CLEF eHealth Evaluation Lab 2013</article-title>
          . In: Forner,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Paredes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (eds.):
          <article-title>Information Access Evaluation</article-title>
          . Multilinguality, Multimodality, and Visualization, LNCC, vol.
          <volume>8138</volume>
          , pp.
          <fpage>212</fpage>
          -
          <lpage>231</lpage>
          . SpringerVerlag, Berlin Heidelberg, Germany (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leroy</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mowery</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
          </string-name>
          , J.:
          <article-title>Overview of the ShARe/CLEF eHealth Evaluation Lab 2014</article-title>
          . In: Kanoulas,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Toms</surname>
          </string-name>
          , E. (eds.):
          <article-title>Information Access Evaluation</article-title>
          . Multilinguality, Multimodality, and Visualization, LNCC, vol.
          <volume>8685</volume>
          , pp.
          <fpage>172</fpage>
          -
          <lpage>191</lpage>
          . Springer-Verlag, Berlin Heidelberg, Germany (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          14.
          <string-name>
            <surname>Poissant</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamblyn</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawasumi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>The impact of electronic health records on time efficiency on physicians and nurses: A systematic review</article-title>
          .
          <source>JAMIA</source>
          <volume>12</volume>
          (
          <issue>5</issue>
          ),
          <fpage>505</fpage>
          -
          <lpage>516</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          15.
          <string-name>
            <surname>Hakes</surname>
            ,
            <given-names>B..</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Whittington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Assessing the impact of an electronic medical record on nurse documentation time</article-title>
          ,
          <source>Journal of Critical Care</source>
          <volume>26</volume>
          (
          <issue>4</issue>
          ),
          <fpage>234</fpage>
          -
          <lpage>241</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          16.
          <string-name>
            <surname>Banner</surname>
          </string-name>
          , L..,
          <string-name>
            <surname>Olney</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Automated clinical documentation: Does it allow nurses more time for patient care? Computers, Informatics, Nursing (CIN) 27(2</article-title>
          ),
          <fpage>75</fpage>
          -
          <lpage>81</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          17.
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapkin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basilakis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dawson</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>A systematic review of speech recognition technology in health care</article-title>
          .
          <source>BMC Medical Informatics., Decision Making</source>
          <volume>14</volume>
          ,
          <issue>94</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          18.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          :
          <article-title>Guidelines for the use of multimedia in instruction</article-title>
          .
          <source>In: Proceedings of the Human Factors., Ergonomics Society 42nd Annual Meeting</source>
          , pp.
          <fpage>1447</fpage>
          -
          <lpage>1451</lpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          34.
          <string-name>
            <surname>Dunn</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Day</surname>
            ,
            <given-names>R. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>K. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coiera</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Learning from hackers: open-source clinical trials</article-title>
          .
          <source>Science Translational Medicine</source>
          <volume>4</volume>
          (
          <issue>132</issue>
          ),
          <year>132cm5</year>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          35.
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nadkarni</surname>
            ,
            <given-names>P. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>D'Avolio</surname>
            ,
            <given-names>L. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uzuner</surname>
            ,
            <given-names>Ö</given-names>
          </string-name>
          :
          <article-title>Overcoming barriers to NLP for clinical text: The role of shared tasks and the need for additional creative solutions</article-title>
          .
          <source>Editorial. JAMIA</source>
          <volume>18</volume>
          (
          <issue>5</issue>
          ),
          <fpage>540</fpage>
          -
          <lpage>543</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          36.
          <string-name>
            <surname>Huang</surname>
          </string-name>
          , C.-C. and
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Community challenges in biomedical text mining over 10 years: success, failure and the future</article-title>
          .
          <source>Briefings in Bioinformatics May</source>
          <volume>1</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          37.
          <string-name>
            <surname>Morita</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kano</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohkuma</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miyabe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aramaki</surname>
          </string-name>
          , E.:
          <article-title>Overview of the NTCIR-10 MedNLP task</article-title>
          .
          <source>In: Proceedings of the 10th NTCIR Conference</source>
          , pp.
          <fpage>696</fpage>
          -
          <lpage>701</lpage>
          . NTCIR, Tokyo, Japan (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          38.
          <string-name>
            <surname>Aramaki</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morita</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kano</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohkuma</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Overview of the NTCIR-11 MedNLP-2 task</article-title>
          .
          <source>In: Proceedings of the 11th NTCIR Conference</source>
          , pp.
          <fpage>147</fpage>
          -
          <lpage>154</lpage>
          . NTCIR, Tokyo, Japan (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          39.
          <string-name>
            <surname>Neamatullah</surname>
            ,,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douglass</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehman</surname>
            ,
            <given-names>L. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reisner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villarroel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>W. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szolovits</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Moody, G. B.,
          <string-name>
            <surname>Mark</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clifford</surname>
          </string-name>
          , G. D.:
          <string-name>
            <surname>Automated</surname>
          </string-name>
          de
          <article-title>-identification of free-text medical records</article-title>
          .
          <source>BMC 8</source>
          ,
          <issue>32</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          40.
          <string-name>
            <surname>Carrell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Malin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aberdeen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bayer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wellner</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Hiding in plain sight: Use of realistic surrogates to reduce exposure of protected health information in clinical text</article-title>
          .
          <source>JAMIA</source>
          <volume>20</volume>
          (
          <issue>2</issue>
          ),
          <fpage>342</fpage>
          -
          <lpage>348</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          41.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehtikunnas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Back</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karsten</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakoski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salanterä</surname>
            ,
            <given-names>S..</given-names>
          </string-name>
          <article-title>Applying language technology to nursing documents: Pros and cons with a focus on ethics</article-title>
          .
          <source>International Journal of Medical Informatics</source>
          <volume>76</volume>
          (
          <issue>S2</issue>
          ),
          <fpage>S293</fpage>
          -
          <lpage>S301</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          42.
          <string-name>
            <surname>Hrynaszkiewicz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Norton</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vickers</surname>
            ,
            <given-names>A. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altman</surname>
            ,
            <given-names>D. G.</given-names>
          </string-name>
          :
          <article-title>Preparing raw clinical data for publication: Guidance for journal editors, authors, and peer reviewers</article-title>
          .
          <source>British Medical Journal (BMC) 340, c181and Trials 11</source>
          ,
          <issue>9</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          43. National Health and Medical Research Council,
          <source>Australian Research Council and Australian Vice-Chancellors' Committee: National Statement on Ethical Conduct in Human Research. National Health Medical Research Council and Australian Vice-Chancellors' Committee</source>
          , Canberra,
          <string-name>
            <given-names>ACT</given-names>
            ,
            <surname>Australia</surname>
          </string-name>
          (
          <year>2007</year>
          updated
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          44.
          <string-name>
            <surname>Jisc</surname>
          </string-name>
          :
          <article-title>The Value and Benefit of Text Mining to UK Further and Higher Education</article-title>
          . Digital Infrastructure. http://www.jisc.ac.uk/whatwedo/programmes/di_directions/strate gicdirections/textmining.aspx (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>