<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Analysis of Variable-Size Vector Based Approach for Formula Searching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pankaj Dadure</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Partha Pakray</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sivaji Bandyopadhyay</string-name>
          <email>sivaji.cse.jug@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering National Institute of Technology Silchar</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The continuously increasing research in the eld of science, engineering, and technology has generated textual and mathematical data in huge amounts. The research in retrieval and searching of textual data achieved the state-of-the-art results while searching and retrieval of mathematical information is in the early stage of research and requires signi cant improvement. Motivated from the concept of formula embedding and term-document matrix, in this paper we have introduced the variable size formula embedding approach where the formula is transformed into the variable size vector. In a vector, each bit represents their occurrence and corresponds to their position in BPIT. The proposed approach has been tested on the Math type data of Math Stack Exchange of ARQMath task and the pro ciency of the same are represented in terms of nDCG´, MAP´ and Precision at 10 measures. The obtained results have shown that the approach of variable size formula embedding requires signi cant improvement to retrieve the syntactically and semantically similar formula.</p>
      </abstract>
      <kwd-group>
        <kwd>Searching</kwd>
        <kwd>Math Stack Exchange</kwd>
        <kwd>Term-Document Matrix</kwd>
        <kwd>Formula Embedding</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Math formulas are the key constitutes in any scienti c documents or
MathBased Question Answering post to communicate and deliver the idea behind it.
Nowadays, the web is a well-known information sharing platform where user's
share their knowledge and opinion. The process of information sharing creates
a huge amount of data which consists of text, images, videos, etc. However,
the data generated from the knowledge sharing sites mainly contains textual
and mathematical information. The retrieval of such information based on the
mathematical formula is a complex and laborious one. Moreover, the number
of textual information retrieval systems have been developed and exhibit their
retrieval potential as well. On the contrary, Mathematical Information Retrieval
(MIR) have faced the many barriers which leads to retrieval of irrelevant
information. For instance, the mathematical symbols owned the prede ned scripting
styles which are qualitatively di erent from linear textual information. In
mathematical information, some mathematical notations posses the alternative
representation like a permutation of k objects chosen from n distinct objects have
several representations: P n, nPk, Pn;k, nPk and P(n, k). In scienti c documents,
k
sometimes, some mathematical notation holds multiple meanings. For instance,
P(x) depicts either the probability of event \x" or multiplication of variable \P"
and \x". In such cases, the context is a signi cant factor to deliver the actual
idea of mathematical notations. In the retrieval of mathematical information,
user query depicts the user need and a well-formed user query may lead to a
more relevant search. The well-formed user query has the potential to retrieve
the relevant search and satisfy the user's needs. To address the above-mentioned
obstacles, several approaches have been implemented and light-up the research
of mathematical information retrieval. Although still there is a huge scope of
improvement in the research eld of mathematical information retrieval.</p>
      <p>The ARQMath lab at CLEF 20201 aims to push-up the state-of-the-art
evaluation design of math-aware information retrieval, and which seek to support the
development and ultimate deployment of new techniques. There are two tasks
namely Answer Finding (Main Task) and Formula Search (Secondary Task) are
organized by ARQMath Lab at CLEF 2020. In answer nding task, the
system returns the ranked list of answers post for the user query selected from the
question post of Math Stack Exchange2 (MSE) whereas the formula search task
returns the ranked list of formulas from MSE which are synthetically or
semantically similar to queried formula. From these two tasks, we are mainly focused on
the formula search task and contributed a variable-size vector based technique.
The data provided by the ARQMath organizer has contained the posts which
hold the textual information and mathematical formulas. As per the exibility
provided by the task organizer, the participant can use both textual and formula
information or only textual information or only formulas. In our proposed
approach, we have used only the formulas extracted from the Math Stack Exchange
site.</p>
      <p>The major contribution of the proposed work are given as follow:
The proposed approach transformed the formula into a variable size vector
of weight.</p>
      <p>Each weight of the vector represents the occurrence of a particular entity in
a formula and correspond to their position in BPIT.</p>
      <p>The proposed approach has been tested on the formulas expressed in
Presentation MathML format which are extracted Math Stack Exchange.</p>
      <p>The structure of the paper is as follows: Section 2 describes related work
in the domain of formula retrieval. Section 3 illustrates the corpus description
and proposed methodology. Section 4 covered the experimental results and their
analysis. Section 5 concludes the paper with a key future direction.</p>
    </sec>
    <sec id="sec-2">
      <title>1 https://www.cs.rit.edu/ dprl/ARQMath</title>
    </sec>
    <sec id="sec-3">
      <title>2 https://math.stackexchange.com</title>
      <sec id="sec-3-1">
        <title>Related Work</title>
        <p>
          The retrieval of mathematical information has seen growing attention in the last
decade and since the mid-1990's the search for Math Formula has been studied.
The past research works have addressed the number of mathematical information
retrieval challenges. However, there is a huge scope for improvement to mitigate
the challenges. In the retrieval of mathematical information, the formula
embedding approach [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] transformed the formula into the vector of size 202 where \0"
depicts the presence and \1" depicts the absence of a particular entity in the
formula. The three layer model of formula representation and searching [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] where
the rst layer choose candidates using spectral matching over tree node pairs,
the second layer aligned a query with candidates and estimates the similarity
based on spectral matching, and the third layer estimates the similarity score of
two representation using linear regression.
        </p>
        <p>
          To furnish the retrieval results for formula query, the Tangent-CFT [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] have
used the two e ective representations of mathematical information i.e.
Symbol Layout Trees (SLTs) and Operator Trees (OPTs) which may contribute to
the more accurate retrieval. In which, symbol layout tree focused on the
appearance of the formula whereas the operator tree focused on the content of
the formula. In which, the tuples have been generated based on the path
between the pair of symbols and embed them using the fast n-gram embedding
model. For state-of-the-art performance, the Tangent-CFT combined the SLTs
and OPTs embeddings and uses the structural similarity for partial-match
formula retrieval. In the research eld of MIR, the transformation of mathematical
information language to natural language is an arduous process. To highlight this
transformation, AnnoMathTex- a recommender system [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] enables the formula
annotation by assigning a meaning to formula identi ers from the surrounded
text.
        </p>
        <p>
          In mathematical information retrieval, relevance measurement is a signi cant
ingredient. For instance, the estimation of the cosine similarity between the
indexed formulas and query formula can also lead to some useful relevant results
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The state-of-the-art results of deep learning techniques in textual information
retrieval motivate the researcher to incorporate it in the retrieval of
mathematical information. The positive impact of LSTM in sequence-to-sequence problem,
LSTM neural network based formula entailment approach [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] formulated the
entailment between the index formulas and query formula. In substructural
matching of formulas, formula retrieval requires the high computation time. To speed
up the formula search in sub-structural similarity, dynamic pruning strategies,
and specialized inverted index produced noticeable outcomes [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. For further
improvement, the query structure representation is associated with posting lists
to boots the overall outcomes of the sub-tree matching. To increase the
accessibility of scienti c documents, the MathAlign system [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] has introduced the
rule-based approach which extracts the latex form of formula and linked the
identi ers of extracted formula to their text description.
        </p>
        <p>
          To analyze the frequency distributions of mathematical expressions in the
large scienti c corpus, Andre et al. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] have incorporated the Zipf's law on a text
corpus. Motivated from the similarity in linguistic properties, they have
introduced a novel approach to rank formulae by their relevance via a customized
version of the ranking function BM25. They have also demonstrated the
applicability of the results by presenting auto completion for math inputs as the rst
contribution to math recommendation systems.
        </p>
        <p>
          Motivated from the emerging association between mathematical formulas
and the textual context in scienti c documents, the topic model called TopicEq
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] has generated the context from a mixture of latent topics, and the equation
has generated by an RNN that depends on the latent topic activation and
enables intelligible processing of equations by considering the relationship between
the mathematical equations and topics. To investigate the feasibility of neural
representation techniques in MIR, the symbol2vec approach [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] learn the vector
representations of formula symbols. For more re ned results, the textual
information has combined with the Formula2vec model and achieves better retrieval
performance. To learn distributed representations of equations, the
unsupervised approach called equation embeddings (EqEmb) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] where the equation has
been treated as singleton word. In which, the semantic representations of
mathematical equations and their surrounding words have embedded and obtained
results are compared with the CBOW, PV-DM, GloVe model where EqEmb-U
achieves the highest performance. The exploratory investigation of the e
ectiveness and use of word embedding [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] where word2vec models have trained on
DLMF and arXiv with slightly di erent approaches for embedding math. The
DLMF trained model discovered the mathematical term similarities and term
analogies and generates the query expansions. The arXiv trained model has
bene cial to extract the textual descriptive phrases for math terms. The word
embedding model mainly focused on term similarity, math analogies, concept
modeling, query expansion, and knowledge extraction.
3
3.1
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Methodology</title>
        <p>Data Description
To evaluate the performance of the proposed formula retrieval approach, the
task organizer provides the dataset collected from the knowledge sharing
platform i.e. Math Stack Exchange (MSE). The dataset comprised the question,
answer, and comment posts. As per the exibility on data, the participant of
the task can use the only formula or only textual or both (formula and textual
information) to perform the formula search. In the dataset, the formulas are
represented in three di erent formats i.e. LATE X, Presentation MathML, and
Content MathML format. These formulas are extracted from the posts of Math
Stack Exchange of the year 2010-2018. The number of formulas comprised in
LATE X, Presentation MathML and Content MathML formats are 28,320,920,
26,075,012 and 25,366,913 of size 1.5 GB, 11.5 GB and 10.9 respectively. Each
format has ve distinct attributes namely formula id, post id, thread id, type,
and formula. To execute the formula search task, we have used only a formula
dataset and out of these three formula representation les, we have used only
the Presentation MathML format. The dataset are available at the o cial site
of ARQMath task3. The metadata about the formula dataset is shown in table
1.
To estimate the performance of the proposed work, the task organizer provided
87 mathematical formulae, each of the formula is selected from the question's
topic of task 1. The topics (queries) for the formula search task provided in an
XML le that has a prede ned format as shown in gure 1. Each topic in XML
le tagged by &lt;topic&gt; and &lt;/topic&gt; tag and each topic has a unique topic
number i.e. B.x where \x" represents the topic number. Formula Id shows the
id of formula, Latex shows the latex representation of formula, Title shows the
question title of the post from which the formula is selected, Question shows the
question body from which the formula is selected and Tags shows the
commaseparated tags of the question.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3 https://www.cs.rit.edu/ dprl/ARQMath/</title>
      <p>
        The system architecture of the proposed approach is shown in gure 2. The
proposed system architecture is inspired by the existing formula embedding
approach [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and term-document matrix [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] where each module work
interdependently to make the faster retrieval and accurate search.
3.3.1 Formulas: The formulas are extracted from the posts of Math Stack
Exchange from 2010-2018 in three di erent format LATE X, Presentation MathML
and Content MathML. To investigate the ability of the proposed approach in
the formula search task, we have used only the Presentation MathML format.
3.3.2 Formula Preprocessing: The prime task of the formula preprocessing
module is to transform the formulas into the uni ed form by removing
irrelevant elements and attributes. In this process, the preprocessing module trim the
tags to their root form for example &lt;mi mathvariant=\\normal""&gt; trimmed
to &lt;mi&gt;, &lt;mo rspace=\\4.2pt""&gt; trimmed to &lt;mo&gt; etc. Some examples of
the preprocessed tags are shown in Table 2. The formula preprocessing
module also discarded a few Presentation MathML tags like &lt;mtext&gt;, &lt;mspace&gt;,
&lt;mstyle&gt; etc. as these tags does not have much contribution regarding semantic
of the mathematical notation.
3.3.3 Formula Embedding: The proposed formula embedding approach
motivates from the existing Bit Position Information Table (BPIT) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and
Term-Document matrix [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] which transformed the formulas into the variable
size vector of weight. Each weight of vector represents the occurrence count of
a particular entity in a formula and correspond to entity position in BPIT. The
process of formula to variable size vector transformation is shown in gure 2. In
the proposed formula embedding approach, entities in a formula are categorized
into the three categories based on the tags. The entities tagged by &lt;mi&gt; comes
under the rst category, entities tagged by &lt;mo&gt; comes under the second
category and the third category holds the essential MathML tags which contribute
to the semantics of formula. As per the position of entities in BPIT, the &lt;mi&gt;
tags hold the positions 0-25, 57-65 &amp; 71-100, the &lt;mo&gt; tags hold the positions
26-45, 66-70 &amp; 101-149 and the positions 46-56 holds the essential MathML tags.
As de ned in the gure 3, the generated vector 3[23, 24, 25], 2[26, 30], 2[47, 49]
from the formula x2 + y2 = z2 where 3, 2 and 2 de ned the occurrence count
of &lt;mi&gt;, &lt;mo&gt; and essential MathML tags respectively and 23, 24, 25, 26, 30,
47, 49 represents the bit position of the &lt;mi&gt;, &lt;mo&gt; and essential MathML
tags in BPIT.
3.3.4 Indexer: In the process of formula searching, the e ective indexing
is the key constituent to speeding up the searching process. After a successful
formula transformation into the variable size vector, the indexer module indexed
the formula vector into an index. Each index stored the three di erent elds
namely embedded formula vector, formula id, and post id from which the formula
is originated.
3.3.5 Query Preprocessing: To test the formula searching e ectivity of the
proposed approach, the ARQMath task organizer provided the queries (Topics)
in LATE X format as described in section 3.2. As we have selected the
Presentation MathML format of formula for formula embedding and to maintain the
uni ed structure between formulas and queries, we have transformed the queries
into the Presentation MathML format using the tool Demo MathType4.
3.3.6 Query Embedding: Query embedding module converts the
preprocessed Presentation MathML query formula into a variable size vector. For query
vector generation, the query embedding module considered the entity position
format of BPIT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In a generated vector, each weight represents the
occurrence of a particular entity in a query formula and their position in BPIT. In
MIR, user's expressed their need in the form of formula and expects that the
system returns the documents/posts which contain syntactically and semantically
similar formulae.
3.3.7 Searcher and Ranker Module: The main objective of the searcher
module is to search for relevant formulas that are syntactically or semantically
similar to query formula and satisfy the user's need. Searching for relevant
information is a time consuming process and requires the e ective formula
representation and indexing technique [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In the proposed architecture, the searcher
module compared the query formula vector with all the indexed formula
vectors and computes the similarity for each indexed formula vector. For similarity
calculation, the proposed approach compared each bit of query vector with the
4 http://www.wiris.com/editor/demo/en/developers
bits of formula vector and those formulas contained the maximum number of
similar bit that formula have maximum similarity score. The process of
similarity calculation is shown in gure 4 where the occurrence count of &lt;mi&gt;, &lt;mo&gt;
and essential MathML tags of formula vector are compared with the occurrence
count of &lt;mi&gt;, &lt;mo&gt; and essential MathML tags of query vector. The reason
behind the consideration of occurrence count is to help in assigning a priority
to those formulae which have a similar number of &lt;mi&gt;, &lt;mo&gt;, and essential
MathML tags.
      </p>
      <p>
        After a successful comparison and similarity calculation between the formula
vectors and query vector, the ranker module retrieves and ranks the post (the
post which contains the search formula) based on the higher similarity score.
Those post of MSE contains more than one similar formula with respect to query
formula, that post of MSE assigned higher priority as compared to those post
which contains the formula only one time. As a nal search result with respect
to query formula, we have retrieved the top 20 posts of MSE which contains the
relevant formulas.
The results of the proposed approach have been estimated using the 45 queries
out of provided 87 queries [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. For each query formula, the proposed approach
retrieves the top 20 posts which contain the relevant formulas and ranked them
based on the maximum similarity score. Those post holds the more than one
similar formulas that post get the higher priority as compared those posts which
contained the formula only once. The obtained search results stored in TSV le
which have six attributes namely query id, formula id, post id, rank, score and
run number where query id represents the query unique id in the form of \B.x"
where x is the query number, formula id represents the unique identi er for the
retrieved formula instance, post id represents the unique identi er of the post
where the formula is contained, rank attributes represents the rank of retrieved
formula, score attribute represents the similarity score between the query formula
and index formula and run number represents the number of runs submitted by
the participant.
      </p>
      <p>
        For evaluating the performance of the proposed approach, trec tool5 is used,
which compared the gold dataset (qrel le) with the result set obtained from
the proposed system. The obtained results of the proposed work for the formula
search task are shown in table 3 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. All the obtained results of the participants
have been ranked based on obtained nDCG´ metric [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. For generic and more
comparative analysis, the results are also estimated in terms of MAP´ and P@10.
The results of the baseline system have achieved the highest score as compared
to all the participant teams. The obtained result disclose that the proposed
approach requires signi cant improvement to generated the state-of-the-art result
for formula search and retrieval tasks.
      </p>
      <p>
        The nDCG´ measure is based on nDCG and is a commonly used scale when
ratings for relevance judgments are available and a single value gure is generated
over a series of ranking lists. The retrieved document receives a gain value (0, 1,
2, or 3) and the rank of each post is gradually decreasing. The discounted gain
values result are obtained and then normalized to [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] by dividing the maximum
discounted cumulative gain possible which is called normalized Discounted
Cumulative Gain (nDCG). The only di erence when nDCG's is calculated is that
unjudged posts are discarded before the measurement is performed. In addition,
the nDCG´ produces a single measure with graded relevance, while Precision@k
and Mean Average Precision (MAP) require binarized judgments in terms of
relevance [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Moreover to nDCG´, the MAP´ is also calculated by removing
unassessed posts and precision 10 posts (P@10).
0.420
0.392
0.135
0.119
0.108
0.100
0.077
0.059
      </p>
    </sec>
    <sec id="sec-5">
      <title>5 https://github.com/usnistgov/trec eval</title>
      <p>The results shown in table 4 discovered that the proposed formula embedding
approach can retrieve the exact match formula with respect to query formula.
The outcomes of (1 + p3i)1=2 retrieved the syntactically similar formula and
which shows the positive impact of the formula embedding in the retrieval of
mathematical information. The proposed formula embedding approach has e
ciently retrieved all those formulae which hold the maximum number of similar
entities and rank them based on the maximum similarity.</p>
      <p>The results shown in table 5 depict that the approach of formula embedding
has the potential to retrieve the partial match formula with respect to query
formula. The obtained result has shown that the proposed approach retrieves
the parent formula or subformula which holds similar entities with respect query
formula.</p>
      <p>The frailty of the formula embedding approach is shown in table 6, which
depicts that the formula embedding approach sometimes retrieves the totally
irrelevant result. The proposed approach e ectively retrieves the relevant formulas
as well but in some cases, it retrieves irrelevant formulas also.
The objective of this paper is to analyze the performance of the formula
embedding approach which transformed the formula into the variable size vector.
Each weight of vector is an occurrence of a particular entity in a formula and
corresponds to their position in BPIT. To estimate relevance between the index
formula vector and query vector, each bit of the query vector is compared with
bits of index formula vector. The feasibility of this approach has been tested on
the Math Stack Exchange corpus of the ARQMath task and the obtained
results revealed the robustness and frailness of the same. The proposed approach
has the ability to retrieves the syntactically similar formula, subformula, or
parent formula. To compared the obtained results, nDCG´ metric is considered as
the primary measure and additionally MAP´ and Precision at 10 measure also
estimated for more generic comparison.</p>
      <p>Our future target is to improve the ranking mechanism by assigning
priorities to entities while calculating the similarities. To reduce the searching and
retrieval time, index optimization is one of our future task which helps to fast the
retrieval and formula searching process. To enrich the e ciency of the formula
search process and retrieval of semantically similar formula, textual data will be
incorporated with the mathematical data.</p>
      <sec id="sec-5-1">
        <title>Acknowledgment</title>
        <p>The authors would like to express their gratitude to the Department of Computer
Science and Engineering and Center for Natural Language Processing, National
Institute of Technology Silchar, India for providing the infrastructural facilities
and support.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alexeeva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharp</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valenzuela-Escarcega</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kadowaki</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pyarelal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morrison</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Mathalign:
          <article-title>Linking formula identi ers to their contextual natural language descriptions</article-title>
          .
          <source>In: Proceedings of The 12th Language Resources and Evaluation Conference</source>
          . pp.
          <volume>2204</volume>
          {
          <issue>2212</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Anandarajan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nolan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Term-document representation</article-title>
          .
          <source>In: Practical Text Analytics</source>
          , pp.
          <volume>61</volume>
          {
          <fpage>73</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dadure</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An empirical analysis on retrieval of math information from the scienti c documents</article-title>
          .
          <source>In: International Conference on Communication and Intelligent Systems</source>
          . pp.
          <volume>301</volume>
          {
          <fpage>308</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Davila</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanibbi</surname>
          </string-name>
          , R.:
          <article-title>Layout and semantics: Combining representations for mathematical formula search</article-title>
          .
          <source>In: Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . pp.
          <volume>1165</volume>
          {
          <issue>1168</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Preliminary exploration of formula embedding for mathematical information retrieval: can mathematical formulae be embedded like a natural language? arXiv preprint</article-title>
          arXiv:
          <volume>1707</volume>
          .05154 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Greiner-Petter</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schubotz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Muller,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Breitinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Cohl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Aizawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Discovering mathematical objects of interest|a study of mathematical notations</article-title>
          .
          <source>In: Proceedings of The Web Conference</source>
          <year>2020</year>
          . pp.
          <volume>1445</volume>
          {
          <issue>1456</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Greiner-Petter</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Youssef</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>B.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schubotz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aizawa</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gipp</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , et al.:
          <article-title>Math-word embedding in math search and semantic extraction</article-title>
          . Scientometrics pp.
          <volume>1</volume>
          {
          <fpage>30</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jansen</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rieh</surname>
          </string-name>
          , S.Y.:
          <article-title>The seventeen theoretical constructs of information searching and information retrieval</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>61</volume>
          (
          <issue>8</issue>
          ),
          <volume>1517</volume>
          {
          <fpage>1534</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Krstovski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.:</given-names>
          </string-name>
          <article-title>Equation embeddings</article-title>
          . arXiv preprint arXiv:
          <year>1803</year>
          .
          <volume>09123</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mansouri</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohatgi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oard</surname>
            ,
            <given-names>D.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanibbi</surname>
          </string-name>
          , R.:
          <article-title>Tangentcft: An embedding model for mathematical formulas</article-title>
          .
          <source>In: Proceedings of the 2019 ACM SIGIR international conference on theory of information retrieval</source>
          . pp.
          <volume>11</volume>
          {
          <issue>18</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pathak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Lstm neural network based math information retrieval</article-title>
          .
          <source>In: 2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>6</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Pathak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pakray</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A formula embedding approach to math information retrieval</article-title>
          .
          <source>Computacion y Sistemas</source>
          <volume>22</volume>
          (
          <issue>3</issue>
          ),
          <volume>819</volume>
          {
          <fpage>833</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sakai</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
          </string-name>
          , N.:
          <article-title>On information retrieval metrics designed for evaluation with incomplete relevance assessments</article-title>
          .
          <source>Information Retrieval</source>
          <volume>11</volume>
          (
          <issue>5</issue>
          ),
          <volume>447</volume>
          {
          <fpage>470</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Scharpf</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mackerracher</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schubotz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Breitinger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gipp</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Annomathtex-a formula identi er annotation recommender system for stem documents</article-title>
          .
          <source>In: Proceedings of the 13th ACM Conference on Recommender Systems</source>
          . pp.
          <volume>532</volume>
          {
          <issue>533</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Yasunaga</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , J.D.:
          <article-title>Topiceq: A joint topic and mathematical equation model for scienti c texts</article-title>
          .
          <source>In: Proceedings of the AAAI Conference on Arti cial Intelligence</source>
          . vol.
          <volume>33</volume>
          , pp.
          <volume>7394</volume>
          {
          <issue>7401</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zanibbi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mansouri</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of arqmath 2020: Clef lab on answer retrieval for questions on math</article-title>
          . In: ARQMath Lab @ CLEF
          <year>2020</year>
          . pp.
          <volume>1</volume>
          {
          <issue>25</issue>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Zhong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohatgi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanibbi</surname>
          </string-name>
          , R.:
          <article-title>Accelerating substructure similarity search for formula retrieval</article-title>
          .
          <source>In: European Conference on Information Retrieval</source>
          . pp.
          <volume>714</volume>
          {
          <fpage>727</fpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>