<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Evaluating a Conceptual Indexing Method by Utilizing WordNet</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Mustapha</forename><surname>Baziz</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">IRIT/SIG</orgName>
								<orgName type="institution">Campus Univ. Toulouse III</orgName>
								<address>
									<addrLine>118 Route de Narbonne</addrLine>
									<postCode>F-31062</postCode>
									<settlement>Toulouse Cedex 4</settlement>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Mohand</forename><surname>Boughanem</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">IRIT/SIG</orgName>
								<orgName type="institution">Campus Univ. Toulouse III</orgName>
								<address>
									<addrLine>118 Route de Narbonne</addrLine>
									<postCode>F-31062</postCode>
									<settlement>Toulouse Cedex 4</settlement>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Nathalie</forename><surname>Aussenac-Gilles</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">IRIT/SIG</orgName>
								<orgName type="institution">Campus Univ. Toulouse III</orgName>
								<address>
									<addrLine>118 Route de Narbonne</addrLine>
									<postCode>F-31062</postCode>
									<settlement>Toulouse Cedex 4</settlement>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Evaluating a Conceptual Indexing Method by Utilizing WordNet</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">E71E717C4B307469CA2F8916A4FBE113</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-25T00:37+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>H3.3 [Information Storage And Retrieval]: Information Search and Retrieval; H.3.1 [Content Analysis and Indexing] -Search process</term>
					<term>Retrieval models Algorithms</term>
					<term>Experimentation Conceptual Indexing</term>
					<term>WordNet</term>
					<term>Documents and Query Expansion</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>This paper describes our participation to the English Girt Task of CLEF 2005 Campaign. A method for conceptual indexing based on WordNet is used. Both documents and queries are mapped onto WordNet. Identified concepts belonging to WordNet synsets are extracted from documents and queries and those having a single sense are expanded. All runs are carried out using a conceptual indexing approach. Results prove a primacy of using queries from the title field of the topics and a slight gain of using stemming compared to the non stemming cases.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>The objective of our participation to the English GIRT task in 2005, was to evaluate the use of a conceptual indexing method based on the WordNet <ref type="bibr" target="#b2">[3]</ref> lexical database. The technique consists in detecting mono and multiword WordNet concepts from both documents and queries and then in using them as a conceptual indexing space. Terms not recognized in WordNet (less than 8%) are also added to complete the representation. Even though they are not useful at the expansion stage, they are used to compare documents and queries at the searching stage. This paper is organized as follows. In section 2, we describe the synoptic scheme of our system which includes the Mercure search engine . In section 3, the tests required for conceptual indexing are formally described: the concept detection and weighting methods in 3.1, and the disambiguation-expansion method in 3.2. Section 4 reports the official evaluation results compared with the median average obtained by all participating systems. Finally, section 5 gives some conclusions and prospects.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Overview of the Approach</head><p>In this section, we describe the conceptual indexing method based on WordNet. The principle involves, being given a document (resp. a query), mapping it onto WordNet and then to extract the concepts (mono and multi terms) that belong to WordNet and appear in the text of the document (resp. the query) <ref type="bibr" target="#b0">[1]</ref>. The extracted concepts are then weighed and marked using part of speech information (POS) to facilitate their expansion. The expansion which we call Short Expansion (or SE) amounts to expanding from the document<ref type="foot" target="#foot_0">1</ref> mono sense WordNet terms (having only one sense) by using all of their synonyms extracted from the synset<ref type="foot" target="#foot_1">2</ref> they belong to, and only one of their hypernym concepts (belonging to their hypernym synset). The indexing method may or may not use expansion and stemming <ref type="bibr" target="#b4">[5]</ref> (according to the run). It includes classical keywords indexing by adding the terms that do not belong to WordNet dictionary. Concepts (mono and multi words) identified in WordNet are marked in the document.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Concepts Detection</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Concept Weighting</head><formula xml:id="formula_0">Document/Query C 1 #n … … C i #n … w 1 _w 2 ,…. …….. w 4_ w 5.. w i …… w 1 w 2 ,… ...w 7 ,… w i ,…. ....w n WordNet Matching C 1 = w 1 _w 2 , C i = w i, …etc.</formula></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Classical Representation</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Conceptual Representation</head><formula xml:id="formula_1">t 1 … … t i …</formula></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Expansion of the concepts using WordNet Synonyms and Hypernyms</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>With or without Stemming WordNet concepts</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Expansion</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Indexing Method</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Run1 Run2 Run3 Run4 Run5</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Single words + stemming</head><p>Figure1-Description of the indexing method used to generate the different runs.</p><p>A total of all, five runs were carried out. They are described in Table2 of section 3.</p><p>In the next section we will explain the main steps of our system: the concept detection and weighting methods used to carry out our experiments.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>I. Detail of the approach</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.1.1">Concepts Detection</head><p>Concept detection consists of extracting mono and multiword concepts from documents and queries that correspond to nodes (synsets) in WordNet. Formally, let consider:</p><formula xml:id="formula_2">D= {w 1 , w 2 , …, w n } (1)</formula><p>the initial document composed of n single words. The result of the concept detection process will be a document D c . It corresponds to:</p><formula xml:id="formula_3">D c = {c 1 , c 2 , …, c m , w' 1 , w' 2,…, w' m' }<label>(2)</label></formula><p>where c 1 , c 2 , , c m are concepts recognized as WordNet entries. These concepts could be mono or multiword. It may also happen that single words w' 1 , w' 2,…, w' m' of the initial document (query) do not belong to the WordNet vocabulary. They will not be used for expanding the document (the query). However, they will be added to the final expanded document in order to be used at the search stage. To detect concepts in the query, we use an ad hoc technique that relies solely on concatenation of adjacent words to identify compound (multiword) concepts in WordNet. In this technique, two alternative ways may be carried on. The first one would be projecting WordNet on the document : all WordNet multiword concepts are mapped onto the document and those occurring in it. This method has the advantage of creating a reusable resource (a document representation made out of WordNet concepts). Its drawback is the possibility to omit concepts which appear in the document and in WordNet under different forms. For example, if WordNet containts the multiword concept "solar battery", a simple comparison with document would miss the same concept appearing in its plural form "solar batteries". The second way, which we adopt in our experiments, follows an the opposite path, projecting the document onto WordNet: for each multiword candidate concept derived by combining adjacent words in the document, we first question WordNet using these words just as they are, and then we use their base forms if necessary.</p><p>Word are combined, as shown in Figure1, according to the longest succession of words for which a concept is detected. In the example of Figure1, the longest concept "chief_operating_officer#n" (#n is used for the POS name) is selected although "chief " and "officer" could also be identified as single word concepts. This concept is defined by WordNet as follow: chief executive officer, CEO, chief operating officer --(the corporate executive responsible for the operations of the firm; reports to a board of directors; may appoint other managers (including a president))</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Example of a document after its projection onto WordNet:</head><p>In Figure <ref type="figure">3</ref> below, we can see a document example from the collection (named GIRT-EN19950120120), after its projection onto WordNet conceptual network. For example health_care_delivery#n is a concept that belongs to a WordNet synset identified in the document. Words that are not tagged (like "ddr" in this example) do not belong to WordNet terminology. Figure3. An example document from the collection after its projection onto WordNet.</p><p>The notations "#n", "#a", "#v", "#r" are used to indicate the part of speech (POS) of the terms belonging to WordNet. They refer respectively to names, adjectives, verbs and adverbs. For the moment, the POS is not used in the index. We need it only to expand the identified mono-sense WordNet terms.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.1.2">WordNet Covering rate for Documents and Queries</head><p>As seen in the previous example, a large majority of the vocabulary used in the collection documents is covered by WordNet. Table1 summarizes the cover rate concerning both queries and documents. More than 92.87% of the vocabulary used in the documents is covered by WordNet and 99.39% (so almost totality!) of the vocabulary used in the queries is covered.</p><p>Concerning compound concepts (or multiterms), they represent about 9% for the document and only 7.83% (0.52 compound term in average) for the queries. Multiterms have often only one sense. It is important to use them in our case, as only mono sense terms from the documents and the queries are expanded in our approach.</p><p>Table1. Statistics on using WordNet to index the English Girt Collection.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>WORDNET TERMS ONLY Total number of TERMS (CLASSICAL) All WN terms WN compounds terms only</head><p>Total no of docs: 151319</p><p>Total no of queries: 25 Documents Queries (1) Documents Queries (1) Documents Queries (1)   Total no of terms  (1) Only Queries using both Title and Description fields (without expansion) are considered in the table.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.1.3">Concepts Weighting</head><p>The extracted concepts (single or multiwords) are then weighted as in the classical keywords case according to a kind of TF.IDF which is also a variant of the OKAPI system <ref type="bibr" target="#b3">[4]</ref>.</p><p>Thus, a weight of a term in a document is given by the following formula [2]: ) , (</p><formula xml:id="formula_4">j i d t Weight i t j d ij j i ij j i tf h d dl h h n N h h tf d t Weight * * 4 3 )) log( * ( * ) , (<label>5</label></formula><formula xml:id="formula_5">2 1 + ∆ + + =<label>(3)</label></formula><p>Where: The objective of this measure is to attenuate the impact of terms having too much high frequency values.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Evaluation</head><p>We submitted five official runs to the monolingual English GIRT task ("GIRT_EN"): CWN_T, C_T, CWN_TD, CWNSE_T and CWNSE_TD. The runs are carried out by using title and/or description fields, using or not the term stemming and by performing or not expansion. They are summarized in Table2. The results obtained by the different runs are summarized in Table3. It should be noticed that an error slipped into the program in the name of query 132 (named by error 232). Consequently, the query 132 is not evaluated at all. The first column of Table <ref type="table" target="#tab_2">3</ref> gives the median average precision (MAP) obtained by our five official runs on all the queries. We give in the second column the same runs when using the query relevance file obtained after submission and with the query 132 corrected. Concerning the official results, as it can be shown in Figure4, the best results are obtained when using only the title field of the topics and stemming the extracted terms (run C_T). Followed by the run CWN_T where WordNet terms are not stemmed, and then the run CWNSE_T where a short expansion (by synonyms and one hypernym) is applied to non polysemic terms. The two last runs (CWN_TD and CWNSE_TD) are obtained when both title and description fields are used to build the queries respectively with and without expansion.</p><p>Figure4. Recall-Precision curves for the fives submitted runs.</p><p>Concerning the non official runs, the results follow the same logic while being better than the official ones. The fourth column of Table <ref type="table" target="#tab_2">3</ref> gives the difference of the global results, for the five runs, between the submitted results and the results obtained after the error has been fixed. Roughly the official results could be enhanced by 10,36% in average for each run by using the query 132 and with changing nothing to the system.</p><p>The reason is that the omitted query (132) brings very good results, which also increases the global result. The detailed results of query 132 are given in Table <ref type="table">4</ref> for the five runs. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Conclusion</head><p>We have evaluated the performances of our conceptual indexing method which consists of matching documents and queries with WordNet. In this method, documents and queries are represented by WordNet nodes. The first remark, when comparing our submitted runs, is that using only title field (runs C_T and CWN_T) from the topics seems to bring the better results than using the title and description fields together. The second remark concerns the use of term stemming. Results showed that stemming indexing terms (run C_T) is slightly better than not stemming them (run CWN_T) when we consider only the first retrieved documents. However, by using a more global judgment (MAP), both cases are close. Another remark concerns the Expansion method used in our experiments. Even though it is made so as to avoid the disambiguation problem (only mono sense terms are expanded), expansion does not seem to bring the best results. The best run is obtained without expansion and by using only the title field of the topics. However, the results obtained by the expansion method, when expanding titles, are better than those obtained when the description fields are used in addition to titles in the queries and without expansion. So we still believe that a more sophisticated expansion method could bring better results <ref type="bibr" target="#b0">[1]</ref>.</p><p>The specificity of the GIRT collection documents could also require some adaptation (to evaluate the usefulness of using hypernymy relation for example).</p><p>Another conclusion concerns the suitability of using WordNet for the domain specific collection. It appears that WordNet largely covers the vocabulary of the English GIRT collection (more than 90% for the documents and practically the entire vocabulary of the 25 used queries) and is suitable to be used for this particular collection.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure2.</head><label></label><figDesc>Figure2. Concept detection method by combining adjacent words.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>ijtf:</head><label></label><figDesc>The frequency of the term in the document ,</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Table2.</head><label></label><figDesc>Description of the official runs. Run Description CWN_T Title field of the topics is used. No stemming is used for WordNet terms. No expansion is used. C_T Title field of the topics is used. Stemming for all terms. No expansion is used. CWN_TD Title and Description fields of the topics are used. No stemming is used for WordNet terms. No expansion is used. CWNSE_T Short Expansion (SE) is used in Queries. Title field of the topics is used. No stemming is used for WordNet terms. CWNSE_TD Short Expansion (SE) is used in Queries. Title and Description fields of the topics are used. No stemming is used for WordNet terms.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0"><head></head><label></label><figDesc></figDesc><graphic coords="7,135.18,105.36,324.96,231.78" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_2"><head>Table 3 .</head><label>3</label><figDesc>Official Results obtained for the five submitted runs compared to the median average.</figDesc><table><row><cell></cell><cell>Median Average</cell><cell>Non official</cell><cell>Increment</cell></row><row><cell></cell><cell>Precision (MAP)</cell><cell>Results</cell><cell>(%)</cell></row><row><cell>CWN_T</cell><cell>0.3411</cell><cell>0.3762</cell><cell>+10,29%</cell></row><row><cell>C_T</cell><cell>0.3411</cell><cell>0.3765</cell><cell>+10,38%</cell></row><row><cell>CWN_TD</cell><cell>0.3223</cell><cell>0.3574</cell><cell>+10,89%</cell></row><row><cell>CWNSE_T</cell><cell>0.3251</cell><cell>0.3579</cell><cell>+10,09%</cell></row><row><cell>CWNSE_TD</cell><cell>0.3235</cell><cell>0.3563</cell><cell>+10,14%</cell></row></table></figure>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0">In the following, the word "document" will refer to both queries and documents in the collection.</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1">WordNet is organised around the notion of Synset (Synonym set). Each Synset contains terms that are synonyms in a given context. Synsets are interrelated by different relations like Hypernymy (Is-a).</note>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">The Use of Ontology for Semantic Representation of documents</title>
		<author>
			<persName><forename type="first">M</forename><surname>Baziz</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Boughanem</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceeding of Semantic Web and Information Retrieval Workshop (SWIR) held in conjunction with the 27 th ACM SIGIR Conference&apos;04</title>
				<meeting>eeding of Semantic Web and Information Retrieval Workshop (SWIR) held in conjunction with the 27 th ACM SIGIR Conference&apos;04<address><addrLine>Sheffield, United Kingdom</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2004">July 25-29, 2004</date>
		</imprint>
	</monogr>
	<note>and Aussenac-Gilles Nathalie</note>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">Mercure at TREC-8&quot; Adhoc, Web, CLIR and Filtering tasks</title>
		<author>
			<persName><forename type="first">M</forename><surname>Boughanem</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Julien</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Mothe</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Soulé-Dupuy</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceeding of Trec-8</title>
				<meeting>eeding of Trec-8</meeting>
		<imprint>
			<date type="published" when="1999">1999</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Wordnet: A lexical database</title>
		<author>
			<persName><forename type="first">G</forename><surname>Miller</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Communication of the ACM</title>
		<imprint>
			<biblScope unit="volume">38</biblScope>
			<biblScope unit="issue">11</biblScope>
			<biblScope unit="page" from="39" to="41" />
			<date type="published" when="1995">1995</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<author>
			<persName><surname>Okapi</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceeding of the 6th International Conference on Text Retrieval TREC</title>
				<editor>
			<persName><forename type="first">D</forename><forename type="middle">K</forename><surname>Harman</surname></persName>
		</editor>
		<meeting>eeding of the 6th International Conference on Text Retrieval TREC</meeting>
		<imprint>
			<date type="published" when="1997">1997</date>
			<biblScope unit="volume">500</biblScope>
			<biblScope unit="page" from="125" to="136" />
		</imprint>
	</monogr>
	<note>TREC-6</note>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">An algorithm for suffix stripping</title>
		<author>
			<persName><forename type="first">M</forename><surname>Porter</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Program</title>
		<imprint>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="130" to="137" />
			<date type="published" when="1980-07">July, 1980</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
