<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Medical Text Indexer System for Indexing Biomedical Literature</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>James G. Mork</string-name>
          <email>mork@nlm.nih.gov</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio J. Jimeno Yepes</string-name>
          <email>antonio.jimeno@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan R. Aronson</string-name>
          <email>alan@nlm.nih.gov</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NICTA Victoria Research Lab</institution>
          ,
          <addr-line>Melbourne</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Library of Medicine</institution>
          ,
          <addr-line>Bethesda, MD</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the face of a growing workload and dwindling resources, the US National Library of Medicine (NLM) created the Indexing Initiative project in the mid-1990s. This cross-library team's mission is to explore indexing methodologies that can help ensure that MEDLINE and other NLM document collections maintain their quality and currency and thereby contribute to NLM's mission of maintaining quality access to the biomedical literature. The NLM Medical Text Indexer (MTI) is the main product of this project and has been providing indexing recommendations based on the Medical Subject Headings (MeSH) vocabulary since 2002. In 2011, NLM expanded MTI's role by designating it as the first-line indexer (MTIFL) for a few journals; today the MTIFL workflow includes about 100 journals and continues to increase. Due to a close collaboration with the Index Section at NLM, MTI continues to grow and expand its ability to provide assistance to the indexers. This paper provides an overview of MTI's functionality, performance, and its evolution over the years.</p>
      </abstract>
      <kwd-group>
        <kwd>Indexing methods</kwd>
        <kwd>Text categorization</kwd>
        <kwd>MeSH</kwd>
        <kwd>MEDLINE</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The NLM Medical Text Indexer (MTI) system [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is the primary product and focus of
the Indexing Initiative [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. MTI produces both semi- and fully-automated indexing
recommendations based on the Medical Subject Headings (MeSH®)1 controlled
vocabulary and has been in use at NLM since 2002. MTI is in daily use to assist
Indexers, Catalogers, and NLM’s History of Medicine Division (HMD) in their indexing
efforts. Every weeknight MTI provides recommendations for approximately 4,000
new citations for Indexing and processes a mixed file of approximately 7,000 old and
new records for both Cataloging and HMD. MTI was also used on a regular basis
between 2002 and 2012 to provide fully-automated keyword indexing for NLM’s
Gateway2 meeting abstract collection, which was not manually indexed. In 2011,
MTI was designated as the First-Line Indexer (MTIFL) for 14 journals (89 in 2013)
1 http://www.nlm.nih.gov/pubs/factsheets/mesh.html
2 http://www.nlm.nih.gov/pubs/factsheets/gateway.html
because of its success with those publications. For MTIFL journals, MTI indexing is
treated like human indexing and, of course, subject to the normal manual review
process. MEDLINE® Indexers and Revisers consult MTI recommendations for
approximately 58% of the articles they index, and the MTI recommendations are tightly
integrated into the Cataloging and HMD system. Although mainly used in indexing
efforts for processing MEDLINE citations3 consisting of identifier, title, and abstract,
MTI is also capable of processing arbitrary biomedical text. MTI provides an ordered
list of MeSH Main Headings (MH), Subheadings (SH), and CheckTags (CT)4 as a
final result. MHs are the main descriptors or headings from the MeSH Vocabulary
(e.g., Lung). SHs are used to qualify the MHs (e.g., Lung/abnormalities means that
the article is about the abnormalities
associated with the Lung more than
the Lung itself), and CTs are a
special type of MHs that are required to
be included for each article and cover
species, sex, human age groups,
historical periods, pregnancy, and
various types of research support (e.g.,
Male).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Processing Overview</title>
      <p>
        The Indexing Initiative explored
several indexing methods [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
eventually implementing two of the best
ones as a prototype indexing system
which became the NLM Medical
Text Indexer (MTI). Normal MTI
processing involves receiving a daily
XML formatted MEDLINE5 file
which contains a list of Completed,
In-Process, and In-Data-Review
citations and a list of Deleted PMIDs
(PubMed® Unique Identifier). All
processing is done offline, and the Fig. 1. MTI Process Flow Diagram
MTI results are then stored in a
database for later use by the Indexers. This preloading of the results is necessary since
MTI takes too long to be done in real time for the Indexers. Fig. 1 depicts the
processing flow as MEDLINE citations are processed through the various components of
the MTI system. Each of the major MTI components is described briefly below.
      </p>
      <sec id="sec-2-1">
        <title>3 http://www.nlm.nih.gov/bsd/mms/medlineelements.html 4 http://www.nlm.nih.gov/mesh/features2003.html 5 http://www.nlm.nih.gov/bsd/licensee/elements_descriptions.html</title>
        <p>
          MetaMap Indexing (MMI) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]: a method that applies a ranking function to concepts
found by MetaMap [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Generally speaking, the MMI ranking function was designed
to indicate the characterizing power or “aboutness” of a given concept for a piece of
text, e.g., a MEDLINE citation. It is the product of a frequency factor and a relevance
factor, which is essentially measured by MeSH Tree depth. For concepts found in the
title of the citation, there is a simplified form of the function which maximizes the
frequency factor.
        </p>
        <p>
          PubMed Related Citations [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]: the neighbors of a document are those documents in
the database that are the most similar to it. The similarity between documents is
measured by the words they have in common, with some adjustment for document
lengths. MTI currently uses two methods for determining PubMed Related Citations
(PRC) for the text it is processing. If MTI is working with a MEDLINE citation and
there are enough indexed PRC defined by the PubMed system6, MTI uses that list of
PRC. If MTI is processing free form text or there is an insufficient number of indexed
PRC, MTI will default to using the in-house TexTool7 implementation of PRC.
MEDLINE is the indexed subset of PubMed.
        </p>
        <p>
          Restrict to MeSH [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]: a method which finds the closest MHs to UMLS®
Metathesaurus®8 concepts. Three basic approaches can be used to map a UMLS concept to
MeSH: through synonyms, through built-in mappings, and through inter-concept
relationships. These approaches can be combined into a strategy that maximizes both
specificity (selected MeSH terms are relevant) and sensitivity (the number of concepts
that fail to be mapped to MeSH is small).
        </p>
        <p>Extract MeSH Descriptors: retrieving the MeSH Heading lines from the PRC in
MEDLINE format and tracking whether the MeSH Heading is a main (starred) term
or not. Note that MTI does not recommend main vs. non-main status to the Indexers,
but the status is tracked internally to see if MTI is improving or not.</p>
        <p>
          Clustering and Ranking [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]: the ranked lists of MHs produced by the methods
described so far must be clustered into a single, final list of recommended indexing
terms. The task here is to provide a weighting of the confidence or strength of belief
in the assignment, and rank the suggested headings appropriately.
        </p>
        <p>
          Post-Processing: once all of the recommendations are ranked and selected, validation
of the recommendations is done based on the targeted end-user. Typically, CTs are
added based on triggers from the text and for the remaining recommended headings, a
machine learning algorithm is applied adding frequently occurring CTs [
          <xref ref-type="bibr" rid="ref8 ref9">8,9</xref>
          ], and then
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>6 http://www.nlm.nih.gov/pubs/factsheets/pubmed.html</title>
        <p>
          7 http://www.ncbi.nlm.nih.gov/CBBresearch/Wilbur/IRET/TexTool/
8 http://www.nlm.nih.gov/pubs/factsheets/umlsmeta.html
finally MTI performs subheading attachment [
          <xref ref-type="bibr" rid="ref10 ref11 ref12">10-12</xref>
          ] to individual headings and for
the text in general.
        </p>
        <p>Not all citations processed by MTI go through all of the components listed above.
MTI has various filtering levels and special handling rules which require different
processing pathways. Basic filtering rules have evolved over time based on
ambiguities in the UMLS Metathesaurus, ambiguity in the text, feedback from Indexers, etc.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>MTI Filtering and Post-Processing</title>
      <p>MTI has three levels of filtering which can be selected depending on the
circumstances. Base Filtering, or High Recall Filtering, is performed for all citations and free
text, regardless of whether any further filtering has been selected or not. High Recall
Filtering is used for MEDLINE indexing recommendations and tends to provide a list
of approximately 25 recommendations with most of the good recommendations near
the top of the list. Balanced Recall/Precision Filtering provides filtering which looks
at the compatibility and context of the recommendations based on what path(s) made
the recommendation and provides a good balance between number of
recommendations and the filtering out of good recommendations. Balanced Recall/Precision
Filtering was developed for use in the fully-automatic processing of the NLM Gateway
abstracts and is now used for MTIFL processing. High Precision Filtering is the last
filtering option and provides the highest level of accuracy by requiring
recommendations to come from both MetaMap (MMI) and PubMed Related Citations (PRC). This
provides a small list of quality MTI recommendations while filtering out many good
recommendations as well. The High Precision Filtering option is not currently used
since it provides such a short list of recommendations.</p>
      <p>Once filtering is accomplished, post-processing is performed regardless of the
filtering level used. Post-processing involves cleaning up the final recommendation list by
removing any terms that survived the filtering process but are invalid for the target
audience, filling out the list of terms by adding CTs, Geographicals, and other MHs
based on the text, a machine learning algorithm, and lookup lists, and then finally
attaching subheadings to the individual MHs and creating a global list of subheadings
applicable to the text.</p>
      <p>
        Since MeSH indexing can be viewed as a categorization task, we use machine
learning in the post-processing stage in an effort to improve both Recall and Precision on
the most frequently used terms in MeSH [
        <xref ref-type="bibr" rid="ref8 ref9">8,9</xref>
        ].
      </p>
      <p>
        MTI’s final step in creating its indexing recommendations is to perform subheading
attachment [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10-12</xref>
        ]. Subheading attachment is currently only done for the Indexers
since Cataloging and HMD do not utilize subheadings. Due to the complexity of the
data manipulation required for subheading attachment, it is not provided as a user
option to MTI. Subheadings are not attached to every MH recommended by MTI; the
subheading attachment algorithms use several linguistic and statistical methods to
determine what is appropriate for each MH based on the text and which subheadings
are allowable for each MH. MeSH specifies a subset of the subheadings that are
allowed for each MH, so the subheading attachment algorithms utilize these rules to
ensure that non-allowed combinations are not recommended by MTI. Based on the
results of two user-centered studies [
        <xref ref-type="bibr" rid="ref13 ref14">13,14</xref>
        ], at most three subheadings are attached to
each MH.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>MTI Performance</title>
      <p>MTI has shown a steady increase in usage and acceptance by the NLM indexers since
2002 when it first started producing recommendations for them. MTI is now a
mature indexing tool that benefits greatly from a close collaborative relationship with its
customers. The strides that MTI has been able to make over the last two years would
not be possible without the continued collaboration with the Index Section providing
much needed expertise and insight to the indexing task.</p>
      <p>MTI was able to provide recommendations for over 93% of the total number of
citations that were indexed in 2012. We use the human indexing as a gold standard and
compare that against the MTI recommendations to calculate Precision, Recall, and
F1measure. Overall F1 has improved from 0.3875 in 2008 to 0.5481 in 2012 (+41.45%).
We look forward to the results of the 2013 BioASQ Challenge to see how MTI
performs against other systems. This will be the first opportunity for such a comparison.</p>
    </sec>
    <sec id="sec-5">
      <title>Future Direction</title>
      <p>Several research topics that are planned for the future include: utilizing full text now
that it is becoming more available, assisting in Gene Link and Chemical Flag
identification, utilizing sections identified in Structured Abstracts to help weight
recommendations, identify whether author/publisher supplied keywords might benefit MTI, and
expanding machine learning usage to help improve problematic MeSH Headings. We
also look forward to expanding the number of MTIFL journals.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The Medical Text Indexer Team benefits from a very close collaboration with the
NLM Index Section. This collaboration provides a deeper understanding of the
manual indexing process and insights into other possible avenues where MTI might be
used to assist in the indexing process at NLM.</p>
      <p>This work was partly supported by the Intramural Research Program of the NIH,
National Library of Medicine. NICTA is funded by the Australian Government as
represented by the Department of Broadband, Communications and the Digital Economy
and the Australian Research Council through the ICT Centre of Excellence program.
We would like to thank our colleagues François Lang and Willie Rogers for providing
direct and indirect support of MTI. We would also like to extend special
acknowledgment to Hua Florence Chang who was the original creator of MTI. Florence’s
foresight has provided us with a robust and tunable program.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aronson</surname>
            <given-names>AR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mork</surname>
            <given-names>JG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gay</surname>
            <given-names>CW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Humphrey</surname>
            <given-names>SM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogers</surname>
            <given-names>WJ</given-names>
          </string-name>
          .
          <article-title>The NLM Indexing Initiative's Medical Text Indexer</article-title>
          . Medinfo.
          <year>2004</year>
          Sept.;
          <year>2004</year>
          :
          <fpage>268</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Aronson</surname>
            <given-names>AR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bodenreider</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            <given-names>HF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Humphrey</surname>
            <given-names>SM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mork</surname>
            <given-names>JG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nelson</surname>
            <given-names>SJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rindflesch</surname>
            <given-names>TC</given-names>
          </string-name>
          , and
          <article-title>Wilbur WJ. The NLM Indexing Initiative</article-title>
          .
          <source>Proc AMIA Symp</source>
          <year>2000</year>
          ;:
          <fpage>17</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Aronson</given-names>
            <surname>AR. The MMI Ranking Function Whitepaper</surname>
          </string-name>
          (
          <year>1997</year>
          ). Available at http://skr.nlm.nih.gov/papers/references/ranking.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Aronson</surname>
            <given-names>AR</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lang</surname>
            <given-names>FM.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>An Overview of MetaMap: Historical Perspective</article-title>
          and
          <string-name>
            <given-names>Recent</given-names>
            <surname>Advances</surname>
          </string-name>
          .
          <source>J Am Med Inform Assoc</source>
          .
          <source>2010 May</source>
          <volume>1</volume>
          ;
          <issue>17</issue>
          (
          <issue>3</issue>
          ):
          <fpage>229</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wilbur</surname>
            ,
            <given-names>W. J.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>PubMed related articles: a probabilistic topic-based model for content similarity</article-title>
          .
          <source>BMC bioinformatics</source>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <fpage>423</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bodenreider</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nelson</surname>
            <given-names>SJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hole</surname>
            <given-names>WT</given-names>
          </string-name>
          ,
          <article-title>and Chang HF</article-title>
          .
          <article-title>Beyond Synonymy: Exploiting the UMLS Semantics in Mapping Vocabularies</article-title>
          .
          <source>Proc AMIA Symp</source>
          <year>1998</year>
          ;:
          <fpage>815</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Medical</given-names>
            <surname>Text Indexer (MTI) Processing Flow Whitepaper</surname>
          </string-name>
          . Available at http://skr.nlm.nih.gov/resource/Medical_Text_Indexer_Processing_Flow.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jimeno-Yepes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mork</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          <article-title>Automatic algorithm selection for MeSH Heading indexing based on meta-learning</article-title>
          .
          <source>International Symposium on Languages in Biology and Medicine</source>
          , Singapore, December,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jimeno-Yepes</surname>
          </string-name>
          , Antonio,
          <string-name>
            <surname>Mork</surname>
            <given-names>JG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronson</surname>
            <given-names>AR</given-names>
          </string-name>
          .
          <article-title>Comparison and combination of several MeSH indexing approaches</article-title>
          .
          <source>AMIA Annual Symposium Proceedings</source>
          . Vol.
          <year>2013</year>
          . American Medical Informatics Association,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Névéol</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mork</surname>
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronson</surname>
            <given-names>A.R..</given-names>
          </string-name>
          <article-title>Automatic Indexing of Specialized Documents: Using Generic vs. Domain-Specific Document Representations</article-title>
          .
          <source>Proc BioNLP 2007 Workshop</source>
          ,
          <fpage>183</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Névéol</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shooshan</surname>
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Humphrey</surname>
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rindflesch</surname>
            <given-names>T.C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Aronson</surname>
            <given-names>A.R.</given-names>
          </string-name>
          <string-name>
            <surname>Multiple</surname>
          </string-name>
          <article-title>Approaches to Fine-Grained Indexing of the Biomedical Literature</article-title>
          .
          <source>Proc Pacific Symposium on Biocomputing</source>
          <year>2007</year>
          ,
          <fpage>292</fpage>
          -
          <lpage>303</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Névéol</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shooshan</surname>
            <given-names>SE</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mork</surname>
            <given-names>JG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronson</surname>
            <given-names>AR</given-names>
          </string-name>
          .
          <article-title>Fine-Grained Indexing of the Biomedical Literature: MeSH Subheading Attachment for a MEDLINE Indexing Tool</article-title>
          .
          <source>AMIA Annu Symp Proc</source>
          .
          <year>2007</year>
          ;:
          <fpage>553</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>A MEDLINE Indexing Experiment Using Terms Suggested by MTI Whitepaper</surname>
          </string-name>
          ,
          <year>June 2002</year>
          . Available at http://ii.nlm.nih.gov/resources/ResultsEvaluationReport.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ruiz</surname>
            <given-names>M.E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Aronson</surname>
            <given-names>A.R.</given-names>
          </string-name>
          <article-title>User-centered Evaluation of the MTI System</article-title>
          ,
          <year>2007</year>
          Whitepaper. Available at http://ii.nlm.nih.gov/resources/MTIEvaluation-Final.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>