<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>OAEI 2020 results for AML and AMLC</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Beatriz Lima</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Faria</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francisco M. Couto</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Isabel F. Cruz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Catia Pesquita</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ADVIS Lab, Department of Computer Science, University of Illinois at Chicago</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>BioData.pt &amp; INESC-ID</institution>
          ,
          <addr-line>Lisboa</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LASIGE, Faculdade de Cieˆncias, Universidade de Lisboa</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>State</institution>
          ,
          <addr-line>Purpose, General Statement</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>AgreementMakerLight (AML) is a scalable and extensible ontology matching system with an alignment repair functionality and a strong focus on the use of external knowledge. In OAEI 2020, AML's development focused mainly on expanding its range of complex matching algorithms, but there were also improvements on its instance matching pipeline and on its ontology parsing algorithm. AML remains the system with the broadest coverage of OAEI tracks, and among the top performing systems overall.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Presentation of the System</title>
      <p>the unique characteristics of the matching tasks in this track and to the unavailability of
the TBox assertions in the HOBBIT datasets.
1.2</p>
      <sec id="sec-2-1">
        <title>Specific Techniques Used</title>
        <p>This section describes only the features of AML that are new for OAEI 2020. It also
describes AMLC, a variant of AML tailored to complex matching. For further
information on AML’s simple matching strategy, please consult AML’s original paper [7] as
well as the AML OAEI results publications of 2016-2018 [4, 3, 5].</p>
        <p>Our main development this year was a modular association rule mining framework
for ontology matching, inspired by the work of Zhou et al. [11]. This strategy
resembles the common market basket analysis, where we take into account how frequently
two entities of different ontologies are related to common instances, given a populated
dataset. Our framework features a central association rule mining algorithm
implementation that selects patterns (i.e., mappings) based on their confidence and support, and a
suite of algorithms devoted to finding individual types of patterns and computing their
confidence and support from among the set of instances. As of the OAEI submission we
had implemented only algorithms for detecting simple class and property mappings, but
we are in the process of implementing algorithms for each type of complex mapping.
1.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Adaptations Made for the Evaluation</title>
        <p>As has been the case in recent OAEI editions, the Link Discovery submission of AML
is adapted to these particular tasks and datasets, as their specificities (namely the
absence of a Tbox) demand a dedicated submission. The same is also true to some extent
of AML’s Complex Matching submission.</p>
        <p>As usual, our submission included precomputed dictionaries with translations, to
circumvent Microsoftr Translator’s query limit.
1.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Link to the System and Parameters File</title>
        <p>AML is an open source ontology matching system and is available through GitHub:
https://github.com/AgreementMakerLight.
2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>2.1</p>
      <sec id="sec-3-1">
        <title>Anatomy</title>
        <p>AML’s OAEI 2020 results are summarized in Table 1 and discussed in the following
subsections.</p>
        <p>AML had a 0.7% increase in precision and a 0.9% decrease in recall, resulting in a
0.2% decrease in F-measure, in comparison with its performance in recent years. These
differences are an unexpected consequence of minor changes in AML’s general
configuration.</p>
        <sec id="sec-3-1-1">
          <title>FLOPO-PTO ENVO-SWEET ANAEETHES-GEMET AGROVOC-NALT</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>Conference Populated Conference Hydrography Geolink</title>
          <p>Populated Geolink
Populated Enslaved
Taxon</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>OntoFarm (ra1-M3)</title>
          <p>OntoFarm (ra2-M3)
OntoFarm (rar2-M3)
OntoFarm (Discrete)
OntoFarm (Continuous)
DBpedia-OntoFarm
HP-MP
DOID-ORDO</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Anatomy (error 0.0)</title>
          <p>Anatomy (error 0.1)
Anatomy (error 0.2)
Anatomy (error 0.3)
Conference (error 0.0)
Conference (error 0.1)
Conference (error 0.2)
Conference (error 0.3)</p>
        </sec>
        <sec id="sec-3-1-5">
          <title>Aggregate (class) Aggregate (property) Aggregate (instance) Aggregate (all)</title>
          <p>FMA-NCI small
AML improved its results on both the FLOPO-PTO and the ENVO-SWEET tasks in
comparison with last year. It was surpassed by two versions of LogMap on the
FLOPOPTO task, but remained the best performing system in the ENVO-SWEET task.
With respect to the new tasks, AML ranked third in the ANAEETHES-GEMET task,
and was the only system able to produce results in the AGROVOC-NALT task.
2.3</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Complex Matching</title>
        <p>AMLC was one of three tools able to generate complex correspondences, and the only
tool able to produce results in the (non-populated) Conference task, which uses the
simple reference alignment as input. While its performance was among the best in most
tasks, it remains mediocre in comparison with its performance in simple matching tasks,
underpinning the fact that there is much room for improvement in complex ontology
matching.</p>
        <p>We unfortunately were unable to finish implementing the suite of pattern mining
algorithms for complex ontology matching in time for this OAEI edition, which likely
would have improved AML’s performance substantially in populated complex tasks.
2.4</p>
      </sec>
      <sec id="sec-3-3">
        <title>Conference</title>
        <p>AML had the exact same results as in recent years, with F1-measures of 74% according
to the full reference alignment (ra1), 70% according to the extended reference alignment
(ra2), 78% according to the discrete uncertain reference alignment, and 77% according
to the continuous one, ranking first in all four evaluation variants. It ranked second in
the evaluation with the violation free version of the extended reference alignment (rar2),
likely because AML’s repair algorithm deliberately does not address conservativity
violations, as we do not subscribe to conservativity as a guiding principle in ontology
matching.</p>
        <p>AML was one of only five systems able to participate in a new unannounced task
consisting in matching the DBpedia to the OntoFarm ontologies, and had the highest
Fmeasure among those five.
2.5</p>
      </sec>
      <sec id="sec-3-4">
        <title>Disease and Phenotype</title>
        <p>AML ranked it third and second in F-measure in the HP-MP and DOID-ORDO tasks,
respectively. However, as has been the trend, AML was one of the systems with the
highest number of unique mappings (i.e., mappings not proposed by any other system).
Since the evaluation in this track is based on a 3-vote consensus alignment, rather than a
true reference alignment, and unique mappings are not otherwise assessed, this severely
affects AML’s evaluation, making its results below average in comparison with other
biomedical matching tasks.</p>
      </sec>
      <sec id="sec-3-5">
        <title>2.6 Interactive Matching</title>
        <p>AML had a lower performance than last year in the Anatomy track, undoubtedly tied to
its change in performance in the non-interactive version of the track. Its results in the
Conference track remained the same. Overall it remains the interactive system that is
the least impacted by the oracle errors.
2.7</p>
      </sec>
      <sec id="sec-3-6">
        <title>Large Biomedical Ontologies</title>
        <p>AML’s performance in this track was similar to last year’s, but with decimal increases in
F-measure across all tasks, likely due to the same changes that affected its performance
in the Anatomy track. It remains the best performing system in five out of the six tasks.
2.8</p>
      </sec>
      <sec id="sec-3-7">
        <title>Knowledge Graph</title>
        <p>Contrarily to last year, AML was able to complete all of the five tasks in a timely
manner, having a global F-measure of 0.85, which ranked it third overall. It had the best
performance in matching classes.
2.9</p>
      </sec>
      <sec id="sec-3-8">
        <title>Link Discovery</title>
        <p>As in previous years, AML and all other participants produced a perfect result (100%
F-measure) in the Spatial track. AML had the highest run time among participating
systems, though this was not true in all tasks.
2.10</p>
      </sec>
      <sec id="sec-3-9">
        <title>Multifarm</title>
        <p>AML’s results were slightly better than last years’, with a 2% increase in F-measure
in the different ontologies modality and a 1% increase in the same ontologies
modality. These differences are due to correcting a minor configuration problem when using
AML’s word-matching algorithm in a multilingual setting.
AML obtained the same results as last year, with an F-measure of 86%, which ranked
it fourth.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>General Comments on the Results</title>
      <p>In 2020, AML was once again the system that tackled the most OAEI tracks and
datasets, and maintained its status as one of best performing and broadest matching
systems competing in the OAEI.</p>
      <p>Nonetheless, there is still some work to be done in terms of complex matching, in order
to be able to provide more robust results. We will strive to refine and improve AML’s
complex matching pipeline, particularly by upgrading our association rule based
approach.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>Like in recent years, AML was the matching system that participated in the most OAEI
tracks and datasets, and it was among the top performing systems in most of them.
AML’s performance was very similar to those of recent years in any of the long-standing
OAEI tracks, as most of our development effort went into tackling new challenges, such
as pattern mining approaches for complex matching.</p>
      <p>Complex matching remains one of the biggest challenges in ontology matching, and
will remain the main focus of AML’s development in the near future.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>DF was funded by the Portuguese FCT Grant 22231 BioData.pt (co-financed by FEDER).
CP and BL are supported by FCT through project SMILAX (PTDC/EEI-ESS/4633/2014),
and the LASIGE Research Unit (UIDB/00408/2020 and UIDP/00408/2020).</p>
      <p>FMC was also funded by PTDC/CCI-BIO/28685/2017. The research of IFC was
partially funded by NSF award III-1618126 and by NIGMS-NIH award R01GM125943.
3. D. Faria, B. S. Balasubramani, V. R. Shivaprabhu, I. Mott, C. Pesquita, F. M. Couto, and
I. F. Cruz. Results of AML in OAEI 2017. In ISWC International Workshop on Ontology
Matching (OM), volume 2032 of CEUR Workshop Proceedings, pages 122–128.
CEURWS.org, 2017.
4. D. Faria, C. Pesquita, B. S. Balasubramani, C. Martins, J. Cardoso, H. Curado, F. M. Couto,
and I. F. Cruz. OAEI 2016 results of AML. In ISWC International Workshop on Ontology
Matching (OM), volume 1766, pages 138–145. CEUR-WS.org, 2016.
5. D. Faria, C. Pesquita, B. S. Balasubramani, T. Tervo, D. Carric¸o, R. Garrilha, F. M. Couto,
and I. F. Cruz. Results of AML Participation in OAEI 2018. In ISWC International Workshop
on Ontology Matching (OM), volume 2288 of CEUR Workshop Proceedings, pages 125–131.</p>
      <p>CEUR-WS.org, 2018.
6. D. Faria, C. Pesquita, E. Santos, I. F. Cruz, and F. M. Couto. Automatic Background
Knowledge Selection for Matching Biomedical Ontologies. PLoS One, 9(11):e111226, 2014.
7. D. Faria, C. Pesquita, E. Santos, M. Palmonari, I. F. Cruz, and F. M. Couto. The
AgreementMakerLight Ontology Matching System. In OTM Conferences - ODBASE, pages 527–541,
2013.
8. C. Pesquita, D. Faria, C. Stroe, E. Santos, I. F. Cruz, and F. M. Couto. What’s in a ”nym”?
Synonyms in Biomedical Ontology Matching. In International Semantic Web Conference
(ISWC), pages 526–541, 2013.
9. E. Santos, D. Faria, C. Pesquita, and F. M. Couto. Ontology Alignment Repair Through</p>
      <p>Modularization and Confidence-based Heuristics. PLoS ONE, 10(12):e0144807, 2015.
10. W. Sunna and I. F. Cruz. In International Conference on GeoSpatial Semantics (GeoS),
pages 82–97. Springer.
11. L. Zhou, M. Cheatham, and P. Hitzler. Towards Association Rule-Based Complex Ontology
Alignment. In X. Wang, F. A. Lisi, G. Xiao, and E. Botoeva, editors, Semantic Technology,
pages 287–303, Cham, 2020. Springer International Publishing.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>I. F.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Palandri Antonelli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Stroe</surname>
          </string-name>
          .
          <source>AgreementMaker: Efficient Matching for Large Real-World Schemas and Ontologies. PVLDB</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <fpage>1586</fpage>
          -
          <lpage>1589</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>I. F.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stroe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Caimi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pesquita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Couto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          .
          <article-title>Using AgreementMaker to Align Ontologies for OAEI 2011</article-title>
          .
          <source>In ISWC International Workshop on Ontology Matching (OM)</source>
          , volume
          <volume>814</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>114</fpage>
          -
          <lpage>121</lpage>
          . CEUR-WS.org,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>