<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Development of Methods for Extracting Information from Pharmacy Line Using Conditional Random Fields</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexey I. Molodchenkov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Artem A. Nikolaev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evgenia A. Mitrokhina</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Federal Research Center “Informatics and Control” of the Russian Academy of Sciences</institution>
          ,
          <addr-line>Moscow</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Moscow Institute of Physics and Technology</institution>
          ,
          <addr-line>Dolgoprudny</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Peoples' Friendship University of Russia</institution>
          ,
          <addr-line>Moscow</addr-line>
        </aff>
      </contrib-group>
      <fpage>340</fpage>
      <lpage>348</lpage>
      <abstract>
        <p>The paper considers the solution to the problem of extracting information from short lines of pharmacological orientation in Russian language. As an example, pharmacy lines are used, from which you need to extract the full name of the drug, manufacturer, form of issue, dosage, number of pieces in a package and some other parameters. To extract this information, a conditional random field (CRF) algorithm was used. There was also created a method for preliminary standardization of the strings to bring string tokens to a single form. More than seven thousand pharmacy lines were marked for the experiments and 2 CRF models were trained - with and without preliminary standardization of the lines. For the model with standardization, the following results were obtained: accuracy for diferent data sets is 0.95 (on the validation set) and 0.89 (on the test set). For the model without standardization, the accuracy is 0.95 (on the validation set) and 0.87 (on the test set).</p>
      </abstract>
      <kwd-group>
        <kwd>Named Entity Recognition</kwd>
        <kwd>Conditional Random Fields</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Extracting information from texts is relevant as it is used to solve a number of
problems. The main goal of the tasks of extracting information from texts is to
convert unstructured text data to some structured form (for example, a table or
a semantic graph) for further processing of the received data.</p>
      <p>Text analysis mainly consists of the following steps:
– vectorization of text;
– application of various methods (for example, machine learning) for their
further processing, depending on the problem being solved.</p>
      <p>
        Text vectorization is converting words to normal form and then converting
them to vector form. For this, methods of tokenization, morphological analysis
and vectorization are used. To convert words to an imperfect form, the libraries
Mystem [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for Russian, pymorphy2 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and nltk [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for Russian and some other
languages can be used. Methods and pre-trained models can be used to vectorize
words and texts, such as: a bag of words [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], word2vec [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], doc2vec [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and others.
      </p>
      <p>At the next stage, depending on the task, regular expressions, rules,
additional dictionaries, machine learning methods, etc. are applied.</p>
      <p>In this paper, we consider the problem of extracting information from short
pharmacy lines containing information about goods sold, for their further
comparison with a predetermined reference book of medicinal products. Such
solutions can be applied in various fields of activity and companies. For example, a
marketing agency can use this information to assess the pharmaceutical market.
Large companies with many warehouses and stores can use this kind of solution
to automate the accounting of their products.</p>
      <p>A feature of the texts used in this work is their small length and high density
of entities that need to be recognized. For example, a pharmacy line contains
information about the name of the drug, manufacturer, batch number, taste, if
available, form of release, dosage, etc. Also, the texts contain many words that
were not previously known (for example, new names of drugs or manufacturers),
a minimum of grammar and many abbreviations. These features severely limit
the application of the most commonly used approaches and algorithms.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem Statement</title>
      <p>The task of extracting information from texts is a Named Entity Recognition
task (NER). A named entity is an n-gram in text for which a class is defined.
The task of recognizing named entities is to select continuous fragments of text
and classify them.</p>
      <p>At the entrance, a pharmacy line in Russian is given approximately of the
following type: "АСКОРБИНОВАЯ К-ТА ГЛЕНВИТОЛ КЛУБНИКА №10
ТАБ.ЖЕВ. КРУТКА". It is necessary to first recognize the name of the drug,
manufacturer, lot number and other parameters in this line, then link them to
the reference name of the drug, manufacturer, lot number, etc. for further search
for this string in the directory.</p>
      <p>Let us list the problems that complicate the solution of this problem, which
are to be solved:
– Abbreviations of some words ("к-та" instead of "кислота").
– Producers recorded in diferent languages ("биодерма лаборатория" and
"BIODERMA LABORATORIES").
– Words that have multiple meanings depending on the context (the word
«мед» as a taste or an abbreviation for the word «медицинский»).
3</p>
      <p>An Overview of Named Entity Recognition Methods
Initially, the NER problem was solved without machine learning at all - using
rule-based systems (for example, regular expressions). This solution stops
working normally as soon as any ambiguities of the natural language come into play,
but even in our task it can be used to determine the batch number, since a
limited number of ways of recording it can be distinguished in the data. This
solution gives us an f1-score of about 0.96 on one dataset and 0.93 on the other.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the authors investigated several diferent ways to recognize names,
dates, locations, phone numbers and times from short messages in Swedish,
including regular expressions. This method shows the best result for dates (0.72
F-measures), the worst - for locations (0.57 F-measures). The paper also shows
that dictionaries and parts of speech significantly improve this result (the average
F-measure increased from 0.65 to 0.84).
      </p>
      <p>
        Progress in solving the NER problem has become the methods of classical
supervised machine learning. In addition, entity dictionaries were actively used,
which did not solve the ambiguity problem, but improved the quality. Among the
algorithms that were actively used then were Support Vector Machine (SVM, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ])
and Conditional Random Fields (CRF, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]), but also decision trees ([
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]),
hidden Markov models ([
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) and others. The disadvantage of these models is that
feature selection is a completely empirical process, primarily based on linguistic
intuition, and then a trial and error method; and the choice of features depends
on the problem, which implies additional research for each new NLP problem. A
more detailed overview of methods for solving the problem of recognizing named
entities can be found in the source [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        If we are talking about modern algorithms, then the problem of recognizing
named entities is solved usually by neural network algorithms using Bi-LSTM
+ CRF (long short-term memory + conditional random fields [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). Pre-trained
embeddings are applied to the Bi-LSTM input, after several layers of Bi-LSTM
and the output is a conditional random field (an undirected graph model, without
which, as a rule, it is impossible to achieve state-of-the-art results). You can also
add capitalization features, parts of speech, morphological features, etc. to the
input to embeddings (Bi-LSTM + CRF + Char + Capitalization + POS).
      </p>
      <p>
        In the article [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the authors tested several variants of neural network
architectures containing char and word Bi-LSTM, CRF, word embeddings, highway
networks, etc. on three Russian-language datasets (Gareev’s dataset, FactRuEval
2016, Persons-1000), and it was the Bi-LSTM + CRF + external word
embeddings model that showed state-of-the-art results (F-measure 87.17, 99.26, 82.10,
respectively).
      </p>
      <p>
        Separately, I would like to mention that short texts difer significantly from
long ones, and standard methods for recognizing named entities will work poorly
for them. This is exactly what is shown in the article [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] - the quality has
dropped from the usual 0.8 - 0.9 to 0.3 - 0.5 for tweets.
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] demonstrates the results of using various existing systems for the task of
recognizing named entities in tweets. Some Twitter-specific methods achieve F1
scores over 0.8, but are still far from the current results achieved with longer news
texts. The authors say that the main reason for the deterioration in results is the
poor use of capital letters (poor capitalization) - this feature is very important
for the task of recognizing named entities. Also, abbreviations and slangs worsen
the quality of words that are not included in the dictionary, but their influence
is no longer so significant.
4
      </p>
      <p>Training CRF Model to Extract Entities from
Pharmacy Strings
The training was carried out on 6000 marked lines, which were combined into a
table. Each row of the table contains the pharmacy line itself, as well as all the
parameters that need to be extracted from it. The piece of the data is in the
table 1. The output is a trained CRF model capable of predicting an ordered
sequence of classes corresponding to these tokens for any ordered sequence of
tokens.
Let’s consider the presented steps of the algorithm in more detail.
4.1</p>
      <p>String Standardization
By standardizing a string in this task, we mean bringing the string tokens to a
single form. The method that standardizes strings does the following conversions:
– Removes extra characters (quotes, brackets, commas)
– Brings tokens in cyrillic to a single form, uses a dictionary of substitutions
for this. At this step, the most frequent errors in the spelling of tokens are
"corrected", the ending is brought to a pre-selected form and abbreviations
are replaced with full words
– In fractions, replaces a comma with a dot
– Removes extra spaces and add spaces where needed.</p>
      <p>Example string before standardization
’ВАКСИГРИП СУСП.В/М И П/К 0,5МЛ/ДОЗА ШПР. №1’
and after it ’ВАКСИГРИП СУСПЕНЗИИ ВНУТРИМЫШЕЧНОГО
ВВЕДЕНИЯ И ПОДКОЖНОГО 0.5 МЛ ДОЗА ШПР №1’</p>
      <p>The application of standardization in this task has several goals:
– This approach allows you to improve the accuracy of the model and learn
better on a small sample (or a smaller sample to achieve similar quality, if
we consider an approach with and without standardization).
– Since we isolate and classify tokens to further search for the closest drug
or product in a directory consisting of all possible options, the second goal
of standardization is to use ordinary equality instead of using metrics to
compare the proximity of tokens. This allows you to use filtering by those
ifelds that are unambiguously standardized in our country.
4.2</p>
      <p>Converting a String to the Form Required to Use CRF
Initially, the data is a table of almost 6,000 labeled rows. Each row of the table
contains the pharmacy row itself, as well as all the parameters that need to be
extracted from it (see Table 1).</p>
      <p>To train the CRF model, it is necessary to present the data in the form of a
table, each row of which contains one token, the number of the pharmacy line
from which this token was taken, as well as the class corresponding to this token
(see Fig. 1).</p>
      <p>Description of possible classes:
– FORM_QN - number of pieces in a package
– FULL_NAME - full name of the drug
– MV - volume
– NM_D - dosage
– NM_F - form of issue
– PROD - manufacturer
– O - does not belong to any of the above classes</p>
      <p>Not all the parameters listed here are required to appear in every line.
As features of the word were used: the word itself in lower case, the last 2
characters of this word, the length of the word and a flag about whether this
token is a number or not. And also the same features for two neighboring tokens.
4.4</p>
      <p>Teaching the CRF Model to Predict the Class for a Word
To train the model and conduct experiments, the entire data set was divided
into training and test samples (the size of the test sample is 20% of the entire
data set).</p>
      <p>The CRF (Conditional Random Fields) method was chosen as a classification
method, because it allows you to independently form a set of features by which
you can vectorize words and texts and is popular for the NER problem, as it
is intended for marking sequences. Using word embedding and other standard
vectorization methods is not suitable for this task. New drugs appear, all words
are specific, and the existing methods and pre-trained models were trained in a
common vocabulary.</p>
      <p>A random field is a multidimensional random variable V, where each
component is a one-dimensional random variable. For convenience, we will assume
that ∀i Vi are discrete and the set of their values is finite. We denote the
implementation of a multidimensional random variable V as v ∈ Ω, where Ω is the
set of all possible configurations. A random field can be represented as a graph,
in which the vertices are the components of the multidimensional random
variable V, the edges are the dependencies between them. A random field is called
Markov if 2 Markovian conditions are satisfied:
1. ∀v ∈ Ω P (V = v) &gt; 0
2. P (Vi = vi|Vj = vj , j ∈ A\ {i}) = P (Vi = vi|Vj = vj , j ∈ δi)
where δi - set of neighbors of the vertex Vi.</p>
      <p>A conditional random field is a Markov random field, in which the set of
random variables is divided into 2 disjoint subsets - X and Y - the set of observable
and hidden variables. The prediction task is to optimally reconstruct the values
of y, provided that we know the observables x. That is, the optimization task is to
maximize the conditional probability p (y | x): y∗ = argmaxyp(y|x). Calculation
of the model p * (y | x) is solved as an optimization problem with given
constraints (the diference between the observation and its estimate must be minimal
and the condition Px p(y|x) = 1 for all x). According to the Hammersley-Cliford
theorem (which connects Markov random fields and the Gibbs distribution), we
need to maximize</p>
      <p>Q ψ c(x,y)
p(y|x) = Pc∈C(QG) ψ c(x,y′ ) ,</p>
      <p>y′ ∈y c∈C(G)
where the factor functions ψ c are usually the exponent of a linear
combination of functions from features with weights that need to be determined during</p>
      <p>K
training ψ c = exp( P fk(xc, yc)θ k). This method belongs to the probabilistic
k=1
methods of classical machine learning. Its implementation has good speed, which
is very important when processing large amounts of information.</p>
      <p>
        More details about the CRF method can be found in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experimental Research</title>
      <p>For the experiments, 2 samples were used. The first sample contains 6,000
pharmacy lines and is randomly divided into training and validation at a ratio of
80%/20%. The second sample is an additional 1000 lines taken from another
dataset, which contains a significant proportion of the unknown drug for the
model, since they were absent in the training sample. This sample was used for
the test.</p>
      <p>The two resulting models (with and without string standardization) were
tested on validation and test datasets. In the tables 2, 3, 4 and 5, you can see
the results of the experiments.</p>
      <p>Vectorization of tokens by n-grams and further comparison of vectors using
cosine distance were used as a baseline. The resulting average accuracy for further
comparison was 0.65.</p>
      <p>The first thing you may notice is better quality of both models compared to
the baseline.</p>
      <p>The model shows the worst results on the test data (especially for MV and
NM_D). This can be explained by the fact that the data in the test set contain
a large number of completely new drugs for the model and have some diferences
from the data on which the training and validation was carried out. For example,
dosages and volumes without specifying units of measurement are more common
in the test set.</p>
      <p>You can also notice that on the validation set string standardization does
not improve the prediction quality, but on the test set, there are noticeable
improvements for volume, dosage and form of release - the classes on which the
standardization method has the most significant influence. The diference with
validation can be explained by the fact that the data in the test set have more
typos and abbreviations that need to be corrected through standardization, so
the consequences of standardization are more noticeable.</p>
      <p>In all experiments the model predicts full name of the drug NM_FULL
best of all, the worst predictable classes are dosage NM_D and volume MV.
Dificulties with dosage and volume may occur because they are too similar and
easy to confuse.
Using the CRF method, it was possible to obtain a model showing good results
in the recognition of named entities in short texts of pharmacological topics.
Accuracy for the validation data is 0.95, for the test data it is 0.89. Deterioration
of results can be explained by the emergence of new drugs that are absent in the
training sample, and by some diferences in the data structure - for example, the
frequent absence of units of measure for volume and dosages. In the future, it is
planned to improve the quality by using combinations of diferent approaches to
build a model for the classification of words and by expanding the set of features
for vectorization of tokens.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. Mystem. https://yandex.ru/dev/mystem/.</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. Pymorphy2. https://pymorphy2.readthedocs.io/en/stable/.</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>3. Natural language toolkit</article-title>
          . https://www.nltk.org/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Harris</given-names>
            <surname>Zellig</surname>
          </string-name>
          .
          <article-title>Distributional structure</article-title>
          .
          <source>Word</source>
          ,
          <volume>10</volume>
          :
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          ,
          <year>1954</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Greg</given-names>
            <surname>Corrado Tomas Mikolov</surname>
          </string-name>
          , Kai Chen and
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Eficient estimation of word representations in vector space</article-title>
          .
          <source>ICLR Workshop Papers</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Baldwin</surname>
            <given-names>T. Lau J. H.</given-names>
          </string-name>
          <article-title>An empirical evaluation of doc2vec with practical insights into document embedding generation</article-title>
          .
          <source>arXiv preprint arXiv:1607.0536</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Tobias</given-names>
            <surname>Ek</surname>
          </string-name>
          , Camilla Kirkegaard, H˚akan Jonsson, and
          <string-name>
            <given-names>Pierre</given-names>
            <surname>Nugues</surname>
          </string-name>
          .
          <article-title>Named entity recognition for short text messages</article-title>
          .
          <source>Procedia - Social and Behavioral Sciences</source>
          ,
          <volume>27</volume>
          :
          <fpage>178</fpage>
          -
          <lpage>187</lpage>
          ,
          <year>2011</year>
          . Computational Linguistics and
          <string-name>
            <given-names>Related</given-names>
            <surname>Fields</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. William S Noble.
          <article-title>What is a support vector machine</article-title>
          ?
          <source>Nature Biotechnology</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Bengong</given-names>
            <surname>Yu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Zhaodi</given-names>
            <surname>Fan</surname>
          </string-name>
          .
          <article-title>A comprehensive review of conditional random fields: variants, hybrids and applications</article-title>
          .
          <source>Artificial Intelligence Review</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Kotsiantis</surname>
          </string-name>
          .
          <article-title>Decision trees: a recent overview</article-title>
          .
          <source>Artificial Intelligence Review</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>L.</given-names>
            <surname>Rabiner</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Juang</surname>
          </string-name>
          .
          <article-title>An introduction to hidden markov models</article-title>
          .
          <source>IEEE ASSP Magazine</source>
          ,
          <volume>3</volume>
          (
          <issue>1</issue>
          ):
          <fpage>4</fpage>
          -
          <lpage>16</lpage>
          ,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. David Nadeau and
          <string-name>
            <given-names>Satoshi</given-names>
            <surname>Sekine</surname>
          </string-name>
          .
          <article-title>A survey of named entity recognition and classification</article-title>
          ,
          <source>Jan</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Changki</surname>
            <given-names>LEE</given-names>
          </string-name>
          .
          <article-title>Lstm-crf models for named entity recognition</article-title>
          .
          <source>IEICE Transactions on Information and Systems</source>
          , E100.D(4):
          <fpage>882</fpage>
          -
          <lpage>887</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. The Anh Le, Mikhail Arkhipov, and
          <string-name>
            <given-names>Mikhail</given-names>
            <surname>Burtsev</surname>
          </string-name>
          .
          <article-title>Application of a hybrid bi-lstm-crf model to the task of russian named entity recognition</article-title>
          . pages
          <fpage>91</fpage>
          -
          <lpage>103</lpage>
          ,
          <year>09 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Alan</surname>
            <given-names>Ritter</given-names>
          </string-name>
          , Sam Clark,
          <string-name>
            <given-names>Oren</given-names>
            <surname>Etzioni</surname>
          </string-name>
          , et al.
          <article-title>Named entity recognition in tweets: an experimental study</article-title>
          .
          <source>In Proceedings of the 2011 conference on empirical methods in natural language processing</source>
          , pages
          <fpage>1524</fpage>
          -
          <lpage>1534</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Leon</surname>
            <given-names>Derczynski</given-names>
          </string-name>
          , Diana Maynard, Giuseppe Rizzo, Marieke van Erp,
          <string-name>
            <surname>Genevieve Gorrell</surname>
            , Rapha¨el Troncy, Johann Petrak, and
            <given-names>Kalina</given-names>
          </string-name>
          <string-name>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <article-title>Analysis of named entity recognition and linking for tweets</article-title>
          .
          <source>Information Processing and Management</source>
          ,
          <volume>51</volume>
          (
          <issue>2</issue>
          ):
          <fpage>32</fpage>
          -
          <lpage>49</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>