<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multilingual</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Asma Damankesh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaspreet Singh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fatima Jahedpari</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khaled Shaalan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Farhad Oroumchian</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Retrieval Information</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Reasoning</institution>
          ,
          <addr-line>Plausible Inference, Information</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper the application of the theory of Human Plausible Reasoning (HPR) has been investigated in the domain of filtering and cross language information retrieval. The theory of Human Plausible Reasoning first has been introduced by Collins and Michalski on early 1990s; it has been applied to IR since 1995. This work is an extension to those experiments which focuses on building a framework for cross language information retrieval. The system built in these experiments utilizes plausible inferences to infer new, unknown knowledge from existing knowledge to retrieve not only documents which are indexed by the query terms but also those which are plausibly relevant.</p>
      </abstract>
      <kwd-group>
        <kwd>Human</kwd>
        <kwd>Plausible</kwd>
        <kwd>Filtering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>ACM Categories and Subject Descriptors: H.3.3 Information Search and Retrieval, Information filtering, Retrieval
models</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>From 1950’s when the first Information Retrieval system has been implemented to date, several theories and
techniques have been introduced and implemented by researchers in IR fields such as different document and query
representation (i.e. Vector Space model, probabilistic models, Language modeling), query expansion and weighing
functions. In all these works, the effort is to simulate the real-life behavior of information seekers and information
finders. For example a reference librarian, although (s)he is not an expert in a particular domain, but can infer what
books or documents could be relevant to an information seeker using general and even superficial knowledge of the
subject. In this work an attempt is made to simulate the reasoning aspect of a reference librarian by modifying the
theory of Human Plausible Reasoning.</p>
      <p>
        Human Plausible reasoning is a relatively new theory for answering questions which is proposed by Collins and
Michalski in 1989. Collins and his collegues have spent 15 years investigating how people can draw conclusions in
an uncertain and incomplete situation by using indirect implications. They have developed a descriptive theory of
human plausible inferences that categorizes the plausible inferences in terms of a set of frequently recurring
inference patterns and a set of transformations on those patterns [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. A transformation is applied on an inference
pattern based on a relationship (i.e. generalization and specialization) to relate available knowledge to the query.
Different experimental implementation of the theory such as adaptive filtering [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], XML retrieval [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or expert
finding [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proves the flexibility and usefulness of HPR in the IR domain.
      </p>
      <p>This research is about creating a framework for multilingual IR where all aspects of retrieval in this environment are
represented as different inferences based on HPR. Our experiments so far has focused on the problem of retrieving
relevant documents by the mean of plausible inferences as well as combining evidences of the relevance. In this
system queries are processed and then represented as single words, phrases, logical terms and logical statements.
Where a logical term represents a relation between two words and/or phrases and a logical statement represents a
relationship between a logical term and one or more single word or/and phrase. Different inferences are applied on
query terms to find documents indexed with these terms. In the next step these terms are transformed into new terms
and their related documents are retrieved. The process of generating new terms from query terms or newly generated
terms could be repeated several times. In this process, some documents will be retrieved through several inferences.
Each of these inferences are considered as an evidence of relevance and their weights are combined together in order
to generate a single weight representing how much a document is relevant. In the context of the Information
Filtering, the documents are user profiles and the queries are documents that are arriving one at a time. The problem
we tried to address in our participation in CLEF this year was to measure the applicability of HPR multilingual
filtering domain and to examine different methods for calculating the certainty and combining evidences of
relevance. The attempt is made to build a framework which is independent of any specific language and can infer
new knowledge which could be in a different language by utilizing relationships and general inferences.
This paper is structured as follow: in the first few sections briefly the theory of Human Plausible Reasoning (HPR)
and the plausible inferences are described. Then the proposed system and inferences are explained. The experiments,
findings and deficiencies of our implementation are explained next. The paper is concluded by providing guidelines
for future research.</p>
    </sec>
    <sec id="sec-3">
      <title>An Introduction to the Theory of Plausible Reasoning</title>
      <p>For 15 years, Collin and his colleagues have been investigating the patterns used by people to reason under
uncertainty and incomplete knowledge. They have concluded that these patterns could be categorized in terms of a
set of frequently reoccurring inference patterns, and a set of transformation on those patterns. An inference applies a
transformation on an inference pattern based on some relationship (i.e. generalization, specialization, similarity,
dissimilarity) to relate available knowledge to the questions.</p>
      <p>
        The theory assumes that a large part of human knowledge is represented in “dynamic hierarchies” that are always
being modified, or expanded. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Concepts are represented by nodes, and are connected to each other by some
relationships. Each node can belong to one or more hierarchy and in each hierarchy it’d viewed from different
prospective. A node can be a clause (Lionas in Figure 1.b), an individual (Lionas in Figure 1.a) or a manifestation of
an individual (Lionas in rainy season).
The primitives of the theory consist of basic expressions, operators and certainty parameters. In the formal notation
of the theory, a statement like “Baghdad is the Capital of Iraq” is written as
capital(Iraq) = {Baghdad},γ = 1.0 where capital is descriptor, Iraq is an argument, Baghdad is a referent
and γ = 1.0 is the certainty parameters that indicates we are 100 % sure that this fact is correct. The pair argument
and descriptor is called logical term. Logical statements are terms associated with one or more referents. Descriptor,
argument and referent could be any node in the hierarchy. In addition to the simple statements, dependencies can
form logical expressions too. Elements of expression in the core theory have been summarized in Figure2. The
theory has many parameters for handling uncertainty but it does not explain how these parameters could be
calculated and this is left for implementations and adaptations. The definition of the most important parameters is
given in Figure 3. The theory provides a rich set of inference transforms that could be applied on one statement to
infer new knowledge from the available ones. Interested reader are referred to references [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ][
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
Baghdad is the capital of Iraq
referent r1, r2,{r2 ...}
argument a1, a2, F(a1)
descritor d1, d2
term d1(a1), d2(a2)
statement d1(a1) = {r1} : γ,ϕ
dependenci es between terms
d1(a1) ↔ d2(a1) :α, β, γ
e.g., Baghdad
e.g., Iraq
e.g., Capital
e.g., capital(Iraq)
e.g., capital(Iraq) = baghdad :1,0.02
e.g., latiude(place) ↔ average_te mp(place) :
      </p>
      <p>moderate, moderate, certain)
(translation :i am certain that latitude constrains average temperatur e with moderate
reliability and that average temperatur e constrains the latitude with moderate reliability.)
implicatio s between statements :
d1(a1) = r1 ⇔ d2(a1) = r2 :α, β, γ e.g. grain(plac e) = {rice...} ⇔ rainfal(place) = heavy :</p>
      <p>high, low, certain
(translation : i am certain that if a place produces rice, it implies the place has heavy rainfals
high reliability. but that if a place has heavy rainfall it only implies the place produces rice with
loa reliability.)
1. Certainty ( γ) : The degree of certainty or belief that an expression is true.
2. Frequency ( φ) : Frequency of the referent in the domain of the descriptor (e.g. a large
percentage of birds fly.)
3. Typicality ( τ) : Degree of typicality of subset withing a set. (e.g. robin is a typical bird and
ostrich is not a typical bird)
4. Dominance( δ) : Dominance of a subset in a set(e.g. chickens are not a large percentage of
birds but a large percentage of barnyard fowl.)
5. Similarity (σ ) : Degree of similarity of one set to another set.
6. Conditiona l Likelihood (α ) : conditiona l likelihood that the right − hand side of a dependency or
implicatio n has a particular value (referent) given that the left − handside has a
particular value.
7. Conditiona l Likelihood (β ) : conditiona l likelihood that the rleft − hand side of a dependency or
implicatio n has a particular value (referent) given that the rightt − hand
side has a particular value.
8. Multiplici ty of the referent( μr) : e.g., many minerals are products by a country like Venezuela) .
9. Multiplici ty of the referent( μa) : e.g., many countries produce a mineral like oil).
10. Acceptabil ity (A) : users feedback</p>
    </sec>
    <sec id="sec-4">
      <title>Proposed System</title>
      <p>Like any other logical system our system has four main elements which are document representation, query
representation, domain knowledge base and a set of inference rules. Here, if the partial description of document can
infer the query, we claim that the document is relevant to the query. Briefly, documents are partially described by
concepts, logical terms and statements, and the knowledge base in created to hold relationships between concepts in
the domain. Inference rules are continuously applied on the query term to expand it to infer other related concepts,
logical terms and statement until a plausibly related document or documents are located.</p>
      <sec id="sec-4-1">
        <title>Document Representation</title>
        <p>In this model, documents are represented partially by a finite set of concepts, phrases, logical terms and statement
that can be directly extracted from the document body or title. The reason why this representation is called partial
document representation is that more terms could be inferred from existing ones that also represent the content of
the documents. Every document identifies its concepts, logical terms and statements by the DOC relation. Three
examples below indicate that the doc#1 is about the concept solider, phrase US_troop, and the logical term capital
(Iraq)
1. DOC(solider) = {doc #1}
2. DOC(US _ troop) = {doc #1}
3. DOC(capital(Iraq)) = {doc #1}</p>
      </sec>
      <sec id="sec-4-2">
        <title>Query Representation</title>
        <p>Each query is processed similar to documents and is partially represented by its concepts, phrases and logical terms
or statement. Each of these concepts, phrases, logical terms and statements are then form the argument of a logical
term where the descriptor is the keyword DOC and the referent is unknown. So a query can be represented as a set
of incomplete logical statements which have the form DOC( partial − desciption) = {?}
Therefore the retrieval process can be viewed as the process of finding referents and completing these incomplete
sentences. Since a document can be retrieved in response to several terms in the query or through several inferences,
therefore the final task is to combine the weights assigned to each document from each inference or term and create
a sorted list of the retrieved documents for each query.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Document Retrieval by Plausible Reasoning</title>
        <p>Like any other information retrieval system, in this system the first step is to find a direct match between the query
representation and a document’s partial description. That is, to locate the document or documents which are indexed
by the query terms. Since the query is always represented as an incomplete statement, the aim of this direct approach
is to complete the statement by finding the referents (documents indexed by the term). This direct approach is
applied on each and every concept, phrase and logical term or statements which could be inferred from the query
terms by applying the inference rule depicted in Figure 4.
DOC(subjec t1) = {doc# }
− − − − − − − − − − − − − − − − − − − − − − − − − − − − −
DOC(subjec t1) = doc#
: γ = F1 (γ1 , A)
: γ1 , A
subject can be a concept (e.g. Iraq), a phrase (US_troop) , a logical term (capital(I raq)) or a
statement (capital(I raq) = Baghdad)
Another case is where the document is indexed by a concept, phrase, logical term or statement which is more
specific or general case of the query term. The theory of plausible reasoning provide us with a rich set of
transformations which could be applied on a concept, phrase, in the descriptor, argument or referent of a logical term
or statement to convert the available statement to another which could be the index term of one or more documents.
In this application of the theory only the GEN and SPEC inference transforms are used to move up and down the
hierarchy. These inference rules are used only to infer new concepts and then the direct approach is applied on the
new concept to retrieve the relevant document(s). Figure 5 illustrates the specialization (SPEC-) based argument
transform by an example. As an example let’s consider the query:
DOC(restaurant(iraq_city)) = {?}
which indicates that there is an interest in documents about cities in Iraq which have restaurants. In the knowledge
base we have the fact that the Baghdad is a city in Iraq. Therefore a new query term restaurant (Baghdad) will be
added to the query representation as a new query term. Once the direct approach is applied on the new term c
document doc#3 is retrieved as a related document to the query.</p>
        <p>DOC(d(a)) = {?}
a'SPECa : δ1, A1
− − − − − − − − − − − − − − − − − − − − − − − − −
d(a') : γ1 = F1(δ1, A1)
apply direct approach :
DOC(d(a')) = {?} : γ1
DOC(d(a')) = {doc#} : δ2, A2
− − − − − − − − − − − − − − − − − − − − − − − − − − −
DOC(d(a)) = doc# : γ = F2(γ1,δ2, A2)</p>
        <p>DOC(restaurant(iraq_city)) = {?}
Baghdad SPEC Iraq _ city : 1.0,1.0
− − − − − − − − − − − − − − − − − − − − − − − − − − −
restaurant(Baghdad) :γ = 1.0
DOC(restaurant(Baghdad)) = {?} :γ = 1.0
DOC(restaurant(Baghdad)) = doc#3 : 0.6,1.0
− − − − − − − − − − − − − − − − − − − − − − − − − − −</p>
        <p>DOC(restaurant(iraq_city)) = doc#3 γ = 0.88
The strength of our belief on the relevance of doc#3 to the query both depends on our belief on the suitability of the
restaurant (Baghdad) as a representation for the query and how well that term is a representative of the content of
the doc#3. Interestingly enough it would not make a difference if for example instead of “Baghdad” we had “دا ”
in our knowledge base. That is why we believe this approach is a general framework that can support multilingual
retrieval.</p>
        <p>A different case is when the document is indexed by a concept which is the referent of a query term. Or when the
document is indexed by a logical term whose referent is a query term. Both cases are illustrated in figure 6 and
figure 7 with examples.
Indirect Approach on Referent</p>
        <p>Indirect Approach on Term
:δ1, A1
:γ = F1(δ1, A1)
DOC(d (a)) = {?}
d (a) = r
r
e.g.</p>
        <p>DOC(restaurant(Baghdad)) = {?}
from KB : restaurant(Baghdad) = Nabil : 0.9,1.0
new conpect :Nabil :γ = 0.94
apply direct approach on the new concept
DOC(Nabil) = {?} :γ = 0.94
DOC(Nabil) = doc#2 : 0.8,1.0
− − − − − − − − − − − − − − − − − − − − − − − − − −
DOC(restaurant(Baghdad)) = doc#2 :γ = 0.84
:δ1, A1, ref _ m
: γ = F3(δ1, A1, ref _ m)
DOC (r) = {?}
d (a) = r
d (a)
e.g.</p>
        <p>DOC (US_troop) = {?}
form KB : force(coll ation) = US_troop : 0.7,1.0,0. 5
new term : force(coll ation) : γ = 0.64
Apply direct app roach on the new term
DOC(force( collation) ) = {?} : γ = 0.64
DOC(force( collation) ) = doc#4 : 0.55,1.0
- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -</p>
        <p>DOC (US_troop) = doc#4 : γ = 0.59
In each one of the above cases, first an inference has been applied to generate a new term and then the direct
inference is used to retrieve the relevant document. The calculation of certainty parameters is discussed below.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>For the experiments on information filtering, first the collection was processed and its single words, phrases and
logical terms and logical statements were extracted. The fact that which term came from which document was
ignored at this stage. In the second step all the phrases were processed and some logical terms were also generated in
this way. For example if we had a phrase such as abc. The following logical terms were generated c(ab) and bc(a).
After creating the knowledge base, the profiles were processed and indexed as documents. Then each document was
read, processed and match against the profiles using plausible inferences. The reasoning was limited to only two
levels of depth because of speed limitations.</p>
      <p>Unfortunately, we were not able to process all the documents and only processed the first xxx documents. This was
due to the fact that this was our first major implementation in Python, and we learned the implementation issues the
hard way! We were not able to run the infile client although we received help from Romaric Besançon and INFILE
team but still we had to run the file version of the client. Because we ran out of time, we were not able to work on
thresholds and refining the output of the system, so our results were generally poor. But we hope we will be able to
do much better next year.</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this work an attempt has been made to adapt the Collins and Mechaslki’s theory of Human Plausible Reasoning
as a multilingual framework for information retrieval and information filtering. In these experiments we were able to
build the relation extractor to build a knowledge base, document processor and query processor based on plausible
inferences. However, due to time limit and implementation issues with Python we were not able to add Arabic to
our knowledge base. Also, we were not able to demonstrate a reasonable performance by the time of the CLEF
deadlines. However, we were able to show how this approach could be used to handle multiple languages.
There are many potential improvements to the current system. First is enriching the knowledge base by
implementing better NLP techniques with the ability to produce more accurate and reliable set of terms and
statements. We need to improve our Arabic text processing to form more logical terms and statements. We need to
experiment with different methods of calculating certainty of inferences and combining evidences because currently
we implemented the simplest methods. Other suggestions are applying context, using heuristic and machine learning
strategies. HPR allows defining context for each one of the relationships in the knowledge base and uses them in
inferences. This improves the quality of inferences which is really needed in an information filtering situation.</p>
    </sec>
    <sec id="sec-7">
      <title>References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>ALLEN</surname>
            <given-names>COLLINS</given-names>
          </string-name>
          , R. MICHALSKl,
          <year>1989</year>
          , “
          <article-title>The Logic Of Plausible Reasoning A Core Theory”</article-title>
          ,
          <string-name>
            <surname>Cognitive</surname>
            <given-names>Science</given-names>
          </string-name>
          , Vol.
          <volume>13</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Oroumchian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Arabi</surname>
          </string-name>
          , E. Ashouri,
          <year>2002</year>
          , “
          <article-title>Using Plausible Inferences and Dempster-Shafer Theory of Evidence for Adaptive Information Filtering</article-title>
          .”,
          <source>4th International Conference on Recent Advances in Soft Computing</source>
          , Nottingham, United Kingdom
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Karimzadehgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Habibi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Oroumchian</surname>
          </string-name>
          ,
          <year>2005</year>
          , “
          <article-title>Logic Based XML Information Retrieval for Determining the Best Element to Retrieve”</article-title>
          , Computer since Springer,
          <fpage>pp88</fpage>
          -
          <lpage>99</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Maryam</given-names>
            <surname>Karimzadehgan</surname>
          </string-name>
          , Geneva G.
          <article-title>Belford, and Farhad Oroumchian, Expert Finding by Means of Plausible Inferences</article-title>
          ,
          <source>International Conference on Information and Knowledge Engineering (IKE'08)</source>
          ,Las Vegas, USA, July
          <volume>14</volume>
          -
          <issue>17</issue>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>ALLEN</surname>
            <given-names>COLLINS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MARK H. BURSTEJN</surname>
          </string-name>
          ,
          <year>1988</year>
          , “
          <article-title>Modeling A Theory Of Human Plausible Reasoning”</article-title>
          ,
          <source>Artificial Intelligence III.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Maryam</surname>
            <given-names>Karimzadegan1</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jafar</surname>
            <given-names>Habibi</given-names>
          </string-name>
          , Farhad Oroumchian,
          <article-title>"XML Document Retrieval by means of Plausible Inferences ", in Advances in XML Information Retrieval, Third Workshop of the INitiative for the Evaluation of XML Retrieval INEX 2004</article-title>
          ,
          <string-name>
            <given-names>Schloss</given-names>
            <surname>Dagstuhl</surname>
          </string-name>
          ,
          <fpage>6</fpage>
          -8
          <source>December 2004, Lecture Notes on Computer Systems</source>
          , , Editors N.
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lalmas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Malik</surname>
            and
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Szlávik</surname>
          </string-name>
          , Springer Verlog LNCS 3493, ISBN:
          <fpage>3</fpage>
          -
          <lpage>540</lpage>
          -26166-
          <issue>4</issue>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Ehsan</given-names>
            <surname>Darrudi</surname>
          </string-name>
          , Masud Rahgozar, Farhad Oroumchian,
          <article-title>"Human Plausible Reasoning for Question Answering Systems"</article-title>
          ,
          <source>International Conference on Advances in Intelligent Systems - Theory</source>
          and
          <article-title>Applications in cooperation with IEEE Computer Society</article-title>
          - AISTA'
          <year>2004</year>
          , Luxembourg,
          <fpage>15</fpage>
          -
          <lpage>18</lpage>
          November
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>