<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main"></title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Fayc</forename><surname>¸al Hamdi</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Bt. G</orgName>
								<orgName type="laboratory">LRI</orgName>
								<orgName type="institution" key="instit1">Universit Paris-Sud</orgName>
								<orgName type="institution" key="instit2">INRIA Futurs</orgName>
								<address>
									<addrLine>2-4 rue Jacques Monod</addrLine>
									<postCode>F-91893</postCode>
									<settlement>Orsay</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author role="corresp">
							<persName><forename type="first">Haïfa</forename><surname>Zargayouna</surname></persName>
							<email>haifa.zargayouna@lipn.univ-paris13.fr</email>
							<affiliation key="aff1">
								<orgName type="laboratory" key="lab1">LIPN</orgName>
								<orgName type="laboratory" key="lab2">UMR 7030</orgName>
								<orgName type="institution" key="instit1">Université Paris 13</orgName>
								<orgName type="institution" key="instit2">CNRS</orgName>
								<address>
									<addrLine>99 av. J.B. Clément</addrLine>
									<postCode>93440</postCode>
									<settlement>Villetaneuse</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Brigitte</forename><surname>Safar</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Bt. G</orgName>
								<orgName type="laboratory">LRI</orgName>
								<orgName type="institution" key="instit1">Universit Paris-Sud</orgName>
								<orgName type="institution" key="instit2">INRIA Futurs</orgName>
								<address>
									<addrLine>2-4 rue Jacques Monod</addrLine>
									<postCode>F-91893</postCode>
									<settlement>Orsay</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Chantal</forename><surname>Reynaud</surname></persName>
							<affiliation key="aff0">
								<orgName type="department">Bt. G</orgName>
								<orgName type="laboratory">LRI</orgName>
								<orgName type="institution" key="instit1">Universit Paris-Sud</orgName>
								<orgName type="institution" key="instit2">INRIA Futurs</orgName>
								<address>
									<addrLine>2-4 rue Jacques Monod</addrLine>
									<postCode>F-91893</postCode>
									<settlement>Orsay</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">3477812C17688E72E831A6DF63EC7EB8</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T22:46+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>TaxoMap is an alignment tool which aim is to discover rich correspondences between concepts. It performs an oriented alignment (from a source to a target ontology) and takes into account labels and sub-class descriptions. Our participation in last year edition of the competition have put the emphasis on certain limits. TaxoMap 2 is a new implementation of TaxoMap that reduces significantly runtime and enables parameterization by specifying the ontology language and different thresholds used to extract different mapping relations. The new implementation stresses on terminological techniques, it takes into account synonymy, and multi-label description of concepts. Special effort was made to handle large-scale ontologies by partitioning input ontologies into modules to align. We conclude the paper by pointing out the necessary improvements that need to be made.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>TaxoMap was designed to retrieve useful alignments for information integration between different sources. The alignment process is then oriented from ontologies that describe external ressources (named source ontology) to the ontology (named target ontology) of a web portal. The target ontology is supposed to be well-structured whereas source ontology can be a flat list of concepts. TaxoMap makes the assumption that most semantic resources are based essentially on classification structures. This assumption is confirmed by large scale ontologies which contain rich lexical information and hierarchical specification without describing specific properties or instances.</p><p>To find mappings in this context, we can only use the following available elements: labels of concepts and hierarchical structures.</p><p>Previous participation of TaxoMap in the alignment contest <ref type="bibr" target="#b1">[2]</ref>, despite positive outcome, have put the emphasis on certain limits:</p><p>-Multi-label concepts: previous version of TaxoMap assumed that a concept has only one label. This leads to loose interesting relations between multi-label concepts. -Large ontologies: TaxoMap were unable to run on real ontologies, such as Agrovoc <ref type="foot" target="#foot_0">3</ref> .</p><p>TaxoMap 2 is a new implementation of TaxoMap which aims to overcome these limits and provides modular code (easily extensible). It introduces a morphosyntactic analysis and new heuristics. Moreover, we propose new methods to partition large ontologies into modules which TaxoMap can handles easily.</p><p>We take part to four tests. Results on benchmarks are almost the same as last year as the philosophy behind TaxoMap reminds the same (oriented alignment, between concepts only). We perform better -in terms of number of mappings generated and runtime-on Anatomy. Library test allows us to perform a new algorithm for ontology partitioning and to experiments our system with a new language (Dutch). Directory test enables to test our alignment tool in real world taxonomy integration scenario.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Presentation of the System</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1">State, Purpose and General Statement</head><p>We consider an ontology as a pair (C, H C ) consisting of a set of concepts C arranged in a subsumption hierarchy H C . A concept c is defined by two elements: a set of labels and subclass relationships. The labels are terms that describe entities in natural language and which can be an expression composed of several words. A subclass relationship establishes links with other concepts.</p><p>Our alignment process is oriented; from a source (O Source ) to a target (O T arget ) ontology. It aims at finding one-to-many mappings between single concepts and establishing three types of relationships, equivalence, subclass and semantically related relationships defined as follows.</p><p>Equivalence relationships An equivalence relationship, isEq, is a link between a concept in O Source and a concept in O T arget with labels assumed to be similar. Semantically related relationships A semantically related relationship, isClose, is a link between concepts that are considered as related but without a specific typing of the relation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Subclass relationships</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Techniques Used</head><p>TaxoMap 2 improves terminology alignment techniques. The use of TreeTagger <ref type="bibr" target="#b2">[3]</ref>, a tool for tagging text with part-of-speech and lemma information, enables to take into account the language, lemma and an use word categories in an efficient way. TaxoMap performs a linguistic similarity measure between labels of concepts. The measure takes into consideration categories of words which compose a label. The words are classified as functional (verbs, adverbs or adjectives) and stop words (articles, pronouns).</p><p>Stop words category enables to ignore these words in similarity computation. Functional words has less power than all the other (noun, etc.). The position of a word in the label is also of importance, a common word between two labels is less important after a preposition than a word that is a head. TreeTagger, however, is error-prone, due essentially to short labels.</p><p>The main methods used to extract mappings between a concept c s in O Source and a concept c t in O T arget are:</p><p>-Label equivalence: An equivalence relationship, isEq, is generated if the similarity between one label of c s and one label of c t is greater than a threshold (Equiv.threshold). -Label inclusion (and its inverse) and hidden inclusion: Then, we consider inclusion of label words: let c t be the concept in O T arget with the highest similarity measure with c s . If one of the labels of c t is included in one of the labels of c s , we propose a subclass relationship &lt; c s isA c t &gt;. Inversely, if one of the labels of c s is included in one of the labels of c t , we propose a semantically related relationships &lt; c s isGeneral c t &gt;. If c t is not the concept with the highest similarity measure, its measure must be greater than a threshold (HiddenInc.thresholdSim) and the highest similarity measure must be greater than another threshold (Hidden-Inc.thresholdMax). The intuition behind this strategy is to extract hidden inclusion. -Reasoning on similarity values : Let c tM ax and c t2 be the two concepts in O T arget with the highest similarity measure with c s , the relative similarity is the ratio of c t2 similarity on c tM ax similarity. If the relative similarity is lower than a threshold (isA.threshold), one of the three following techniques can be used:</p><p>• the relationship &lt; c s isClose c tM ax &gt; is generated if the similarity of c tM ax is greater than a threshold (isCloseBefore.thresholdMax) and if one of the labels of c s is included in one of the labels of c tM ax . • the relationship &lt; c s isClose c tM ax &gt; is generated if the similarity of c tM ax is greater than a threshold (isClose.thresholdMax). • an isA relationship is generated between c s and the father of c tM ax if the similarity of c tM ax is greater than a second threshold (isA.thresholdMax). -Reasoning on structure: an isA relationship is generated if the three concepts in O T arget with the highest similarity measure with c s have similarity greater than a threshold (Struct.threshold), and has a common father.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3">Partitioning of large scale ontologies</head><p>We propose two methods of ontology partitioning. The aim of our methods is to have minimum blocs to align with maximal number of concepts (that TaxoMap is able to handle). The originality of our methods is that they are alignment oriented, that means that the partitioning process is influenced by the mapping process.</p><p>The two methods relies on the implementation of PBM <ref type="bibr" target="#b3">[4]</ref> algorithm. PBM partitions large ontologies into small blocks (or modules) and construct mappings between the blocks, using predefined matched class pairs, called anchors to identify related blocks. We only reuse the partitioning part and the idea of anchors, but adapt them in order to take into account the alignment process in the partitionning. We identify the set of anchors as the set of concepts which have the same label in the two ontologies. Even on very large ontologies, this set is computable with a fast and strict equality measure. We also used the possible dissymmetry between ontologies to order the partitionning: if one ontology is well-structured, it will be easier to split it up into cohesive modules, and its partitionning can be used as guideline to partition the other ontology.</p><p>The methods proposed are as follows:</p><p>-Method1 (see figure <ref type="figure" target="#fig_1">1</ref>  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.4">Adaptations made for the Evaluation</head><p>We do not make any specific adaptation in the OAEI 2008 campaign. All the alignments outputted by TaxoMap are uniformly based on the same parameters. For library test, the language was set to nl (for Dutch). We had, however, fixed confidence values depending on relation types.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.5">Link to the system and parameters file</head><p>TaxoMap requires :</p><p>-Mysql -Java (from 1.5) -TreeTagger 4 with its language parameter files.</p><p>The version of TaxoMap (with parameter files) used in 2008 contest can be downloaded from:</p><p>-http://www.lri.fr/˜haifa/TaxoMap.jar: a parameter lg has to be specified it denotes the language of the ontology. For example TaxoMap.jar fr to perform alignment on ontologies in French. If no language is specified, it is supposed to be English. -http://www.lri.fr/˜haifa/TaxoMap.properties: a parameter file which specifies:</p><p>• The command to launch tree-tagger.</p><p>• Treetagger word categories that has to be considered as functional, stop words and prepositions. • The RDF output file.</p><p>• Different thresholds of similarity, depending on the method used.</p><p>-http://www.lri.fr/˜haifa/dbproperties.properties: a parameter file which contains the user and password to access to MySql.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.6">Link to the Set of Provided Alignments</head><p>The alignments produced by TaxoMap are available at the following URLs: http://www.lri.fr/˜haifa/benchmarks/ http://www.lri.fr/˜haifa/anatomy/ http://www.lri.fr/˜haifa/directory/ http://www.lri.fr/˜haifa/library/ 3 Results</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">Benchmark Tests</head><p>Since our algorithm only considers labels and hierarchical relations and only provides mapping for concepts, the recall is low even for the reference alignment. The overall results are almost similar -with no surprise-to those of last year.</p><p>The whole process of alignment costs less than 2 minutes. 4 http://www.ims.uni-stuttgart.de/projekte/corplex/TreeTagger/</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Anatomy Test</head><p>The anatomy real world case is to match the Adult Mouse Anatomy (denoted by Mouse) and the NCI Thesaurus describing the human anatomy (tagged as Human). Mouse has 2,744 classes, while Human has 3,044 classes. As last year, we considered Human as the target ontology as is it well structured and larger than Mouse.</p><p>TaxoMap gains considerably on runtime, it performs the alignment (with no need to partition) in about 25 minutes which is better than last year where TaxoMap took about 5 hours to align the two ontologies.</p><p>TaxoMap generates much more mappings than last year. Only about 200 concepts were left unmapped, whereas last year it was nearly 900. As only equivalence relationships will be evaluated, we change different mapping types to equivalence with these confidence values:</p><p>-(type1) For isEq and isClose relations, confidence value was set to 1.</p><p>-(type2) For isA relations generated by label inclusion, confidence value was set to 0.8. -(type3) For isA relations generated by structural technique or by relative similarity method, confidence value was set to 0.5.</p><p>TaxoMap discovers 2 533 mappings: 1 208 type1 relations, 1 190 type2 relations and 135 type3 relations. The improvement in comparaison with last year results relies on the use of TreeTagger and on taking into account synonymy.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3">Directory Test</head><p>The directory task consists of Web sites directories like Google, Yahoo! or Looksmart. To date, it includes 4,639 tests represented by pairs of OWL ontologies. TaxoMap takes about 40 minutes to complete all the tests.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4">Library Test</head><p>The library task includes two SKOS thesauri GTT and Brinkman thesauri. Since Tax-oMap focuses on Web ontologies expressed in RDFS and OWL, we have to adopt two OWL version ontologies transformed by campaign organizers in this task. GTT owns 35,000 classes, while Brinkman thesauri owns 5,000 classes. The main drawback of using OWL ontologies is that there is no distinction in OWL descriptions (rdfs:label statements) between skos:prefLabel, skos:altLabel and skos:hiddenLabel statements, which removes the subtle distinctions that exist between these different properties.</p><p>We applied the first method of partitioning, this is due to the fact that only 3535 anchors were discovered and that the two ontologies were poorly structured. As the method2 relies on these two informations simultaneously, the partitioning results were not judged relevant.</p><p>The partitioning of Brinkman thesauri (considered as target ontology) leads to 227 modules, the largest module contains 703 concepts. GTT (source ontology) is partitionned into 18 306 modules, 16 265 modules contain only one concept, the largest module contains 517 concepts. We performed 212 combinations that leads to 3 217 mappings.</p><p>The fact that the total number of mappings is less than the number of found anchors is due to the fact that anchors are computed between labels (a concept described by three labels can have three anchors, which is not the case for mappings, where a concept is matched to only one concept). As alignments are performed between modules, this can lead to loose some potential mappings. This is particularly the case of all modules that contain only one concept, as they are ignored by the alignment process.</p><p>As skos relations will be evaluated, we change different mapping types to skos ones with these confidence values:</p><p>-(type1) isEq relations become skos:exactMatch with a confidence value set to 1.</p><p>-(type2) isA relations become skos:narrowMatch with a confidence value set to 1 for label inclusion, 0.5 for relations generated by structural technique or by relative similarity method. -(type3) isGeneral relations become skos:broadMatch with a confidence value set to 1. -(type4) isClose relations become skos:relatedMatch with a confidence value set to 1.</p><p>Generated mappings are as follows: 1 872 type1 relations, 1 031 type2 relations, 274 type3 relations and 40 type4 relations. The whole process of alignment costs about 40 minutes. The partitioning process costs nearly 2 hours. The language of both thesauri is Dutch, we launched tree-tagger with Dutch parameter file. The main difficulty is that there were no Tagset description given for this language and it was difficult to specify word categories needed for the linguistic similarity method.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">General Comments</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1">Results</head><p>TaxoMap 2 significantly improves the results on the previous version of TaxoMap in terms of runtime and number of generated mappings. The new implementation offers extensibility and modularity of code. TaxoMap can be parameterized by the language used in ontologies and different thresholds. We put the emphasis on terminological alignment by taking into account synonymy and multi-label concepts. Our partitioning algorithms allows us to participate to tests with large ontologies.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2">Future Improvements</head><p>The following improvements can be made to obtain better results:</p><p>-Use of WordNet as a dictionary of synonymy. The synsets can enrich the terminological alignment process if an a priori disambiguation is made. -To develop the remaining structural techniques which proved to be efficient in last experiments <ref type="bibr">[5] [6]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">Conclusion</head><p>This paper reports our participation to OAEI campaign with a new implementation of TaxoMap. Our algorithm proposes an oriented mapping between concepts. TaxoMap 2 is better now than last year. Due to partitioning, it is able to perform alignment on realworld ontologies. Our participation in the campaign allows us to test the robustness of TaxoMap, our partitioning algorithms and new terminological heuristics.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head></head><label></label><figDesc>Subclass relationships are usual isA class links. When a concept c S of O Source is linked to a concept c T of O T arget with such a relationship, c T is considered as a super concept of c S .</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>): 1 .</head><label>1</label><figDesc>Use PBM algorithm to partition the target ontology O T into some blocs B T i . 2. Identify the set of anchors included in each module B T i . This set will be the kernel or center CB Si of the future module B Si which will be generated from the source ontology O S . 3. Use PBM algorithm to partition the source ontology around the identified centers CB Si . 4. Align each module B Si with the corresponding module B T i . -Method2 (see figure 2): 1. Partition the target ontology O T by modifying PBM algorithm in order to take into account anchors. Generated modules contain coherent set of concepts that maximize the number of anchors. 2. Partition the source ontology O S the same way then step 1. The interesting anchors that influence partitioning are those that goes in the same module generated from O T . 3. Align modules that share maximal number of anchors.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Fig. 1 . 2 .</head><label>12</label><figDesc>Fig. 1. Method1 for partitioning Fig. 2. Method2 for partitioning</figDesc><graphic coords="4,134.77,449.46,165.85,97.22" type="bitmap" /></figure>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_0">http://www4.fao.org/agrovoc/</note>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">An Information-Theoretic Definition of Similarity</title>
		<author>
			<persName><forename type="first">D</forename><surname>Lin</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">ICML. Madison</title>
				<imprint>
			<date type="published" when="1998">1998</date>
			<biblScope unit="page" from="296" to="304" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">TaxoMap in the OAEI 2007 alignment contest</title>
		<author>
			<persName><forename type="first">H</forename><surname>Zargayouna</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Safar</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Reynaud</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proceedings of the ISWC&apos;07 workshop on Ontology Matching OM</title>
				<meeting>the ISWC&apos;07 workshop on Ontology Matching OM</meeting>
		<imprint>
			<date type="published" when="2007">2007</date>
			<biblScope unit="volume">07</biblScope>
			<biblScope unit="page" from="268" to="275" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Probabilistic Part-of-Speech Tagging Using Decision Trees</title>
		<author>
			<persName><forename type="first">H</forename><surname>Schmid</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">International Conference on New Methods in Language Processing</title>
				<imprint>
			<date type="published" when="1994">1994</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Partition-based block matching of large class hierarchies</title>
		<author>
			<persName><forename type="first">W</forename><surname>Hu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Zhao</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Qu</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. of the 1st Asian Semantic Web Conference (ASWC06)</title>
				<meeting>of the 1st Asian Semantic Web Conference (ASWC06)</meeting>
		<imprint>
			<date type="published" when="2006">2006</date>
			<biblScope unit="page" from="72" to="83" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">When usual structural alignment techniques don&apos;t apply The</title>
		<author>
			<persName><forename type="first">C</forename><surname>Reynaud</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Safar</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">ISWC&apos;06 workshop on Ontology matching</title>
				<imprint>
			<date type="published" when="2006">2006</date>
		</imprint>
	</monogr>
	<note>OM-06</note>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Exploiting WordNet as Background Knowledge The</title>
		<author>
			<persName><forename type="first">C</forename><surname>Reynaud</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Safar</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">ISWC&apos;07 Ontology Matching</title>
				<meeting><address><addrLine>OM-</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2007">2007</date>
			<biblScope unit="volume">07</biblScope>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
