<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Improving Multi-label Classification by Means of Cross-Ontology Association Rules</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
				<date type="published" when="2015-10">October 2015</date>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Fernando</forename><surname>Benites</surname></persName>
							<email>fernando.benites@uni-konstanz.de</email>
							<affiliation key="aff0">
								<orgName type="department">Department of Computer and Information Science</orgName>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Elena</forename><surname>Sapozhnikova</surname></persName>
							<email>elena.sapozhnikova@uni-konstanz.de</email>
							<affiliation key="aff0">
								<orgName type="department">Department of Computer and Information Science</orgName>
							</affiliation>
						</author>
						<title level="a" type="main">Improving Multi-label Classification by Means of Cross-Ontology Association Rules</title>
					</analytic>
					<monogr>
						<imprint>
							<date type="published" when="2015-10">October 2015</date>
						</imprint>
					</monogr>
					<idno type="MD5">E43E9C9BF6B9947A2898D32C805BA52E</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T23:37+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Recently several methods were proposed for the improvement of multi-label classification performance by using constraints on labels. Such constraints are based on dependencies between classes often present in multi-label data and can be mined as association rules from training data. The rules are then applied in a post-processing step to correct the classifier predictions. Due to properties of association rule mining these improvement methods often achieve low improvement expressed mostly in the better prediction of large classes. In the presence of class ontologies this is undesirable because larger classes correspond to higher levels in hierarchies presenting general concepts and can thus be trivial. In this paper we overcome the problem by focusing on improving multi-label classification performance on small classes. We present a new method of improvement based on mining cross-ontology association rules which is best suited for classification with multiple class ontologies, but can also be applied to multi-label classification with one class taxonomy.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>The increasing popularity of ontologies in different areas has led to the availability of data that can be annotated with multiple classes coming from different class taxonomies. This is a special case of multi-label classification. Generally, combining information from the ontologies providing different insights into a domain can be helpful in discovering new cross-ontology associations not evident from only one ontology. For example, if a film is classified by its genre in a genre ontology and by the producing company in an ontology of producers, one can find a possible interesting relation between a certain genre and a producing company, specialized in this genre. Recently, data mining techniques such as association analysis were applied to finding valuable cross-ontology Association Rules (ARs) between multiple ontologies corresponding to distinct categorizations of genes in bioinformatics <ref type="bibr" target="#b1">[2,</ref><ref type="bibr" target="#b8">9]</ref>.</p><p>On the other hand, useful information from such cross-ontology rules can be successfully employed to improve performance in multi-label classification: ARs found among classes of multiple ontologies can be used to correct predicted labels because the presence of a certain class or classes can be helpful for predicting another one. For example, a proper application of the association between a certain genre and a producing company specializing in this genre, as discussed above, can increase the probability of correctly predicting a genre, providing the corresponding company has been already correctly predicted. Thus it should lead to an improvement in classification performance. This is especially important with respect to very large ontologies with many thousands of classes, which are usually difficult to deal with.</p><p>Another problem with class ontologies that has not yet been dealt with sufficiently in recent research is that classifier performance in such a case is largely dominated by more general classes higher in the hierarchy because they are more present in the data and hence simpler to predict for a classifier. On the other hand, such general classes do not often provide interesting information and are sometimes trivial. For this reason, in mining cross-ontology rules rare association rules <ref type="bibr" target="#b13">[14]</ref> are preferred, especially in large ontologies <ref type="bibr" target="#b1">[2]</ref>. Similarly, in the improvement of multi-label classification performance, prediction improvement is more interesting for small and more specific classes in comparison to larger ones. As existing methods have not yet addressed this problem, our paper will focus on mining rare cross-ontology ARs and applying them to the multi-label classification improvement on small classes. For this purpose, a special interestingness measure well-suited for mining rare rules is utilized.</p><p>An important difference between our approach and several state-of-the-art methods discussed in the next section is that they use constraints for labels of one labelset and not two different class ontologies. Further, our approach focuses on rare labels, i.e. the ones with low support, since they are normally the greater part of the labelset.</p><p>The rest of the paper is organized as follows. A brief overview of the approaches to multi-label improvement with ARs is given in Section 2. Afterwards, our approach is explained in Section 3, followed by the experiments of Section 4. In Section 5 we conclude the paper.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Related Work</head><p>Recently several approaches to improving multi-label classification performance were proposed that dealt with dependencies between classes present in multilabel data. Some of them belong to the field of multi-label classification with constraints and apply constraints on labels to performance improvement usually in a post-processing step. The constraints are often mined in form of ARs from training data <ref type="bibr" target="#b7">[8,</ref><ref type="bibr" target="#b9">10]</ref>.</p><p>The initial work was devoted to prediction corrections within the ranking by pairwise comparison framework <ref type="bibr" target="#b9">[10]</ref>. The constraints were in the form of manyto-one ARs labelset→label , i.e. implications from a labelset to a single label. They could be positive or negative which involves either setting or removing a consequent label as a result of the presence of an antecedent label combina-tion. The constraint rules were extracted using a standard support-confidence AR mining framework in order to change the predicted rankings. The results of the method obtained on real-world datasets were negative: no improvement in comparison to the baseline performance was observed. For this reason, this method will not be used for comparison with the proposed method below.</p><p>A more recent approach of <ref type="bibr" target="#b7">[8]</ref> used only one-to-one ARs label i → label j in order to improve SVM performance in the Binary Relevance (BR) setting. The rules to apply were chosen by minimizing the Ranking Loss performance measure through a cross-validation process on the training data. The selected rules were then applied to predicted label rankings in the test phase, if an antecedent label was set, boosting the score of the corresponding consequent label. The improved results were obtained for two real-world datasets Yeast and Reuters. AR mining was based on the standard support-confidence framework. A subsequently extended approach <ref type="bibr" target="#b5">[6]</ref> differs in that it uses subsets of labels gathered by clustering, and also extracts negative and many-to-one ARs. Still the evaluation of the extended approach was restricted to smaller datasets than in the earlier paper (e.g. the Reuters dataset was not included), perhaps pointing to a higher complexity of the algorithm, making it probably inapt to be applied to large datasets. Taking this into account, we selected only the initial method of <ref type="bibr" target="#b7">[8]</ref> for comparison and will refer to it as Label Constraints for SVMs (LCS).</p><p>In contrast to the discussed post-processing methods evaluated in a certain multi-label classification setting (either pairwise or BR), a more general approach, Label Reduction with Association Rules (LRwAR), was proposed in <ref type="bibr" target="#b4">[5]</ref>. It includes pre-and post-processing for the reduction of the label dimensionality. First, ARs are extracted and those labels that are only in the consequents of rules are removed from the data to be learned. Then a multi-label classifier is applied to the classification problem with a reduced labelset. After classification, the rules are applied to recover missing labels. An advantage of this approach is a shorter time needed to train a classifier on a smaller labelset. It also used the standard support-confidence framework to mine ARs, although its recent extension <ref type="bibr" target="#b3">[4]</ref> proposed Conviction instead of Confidence. However it was not shown to provide significantly better results. The base method was evaluated on different multi-label classifiers including ML-kNN, BP-MLL and C4.5 (the latter in BR and label powerset settings) as well as six datasets. On several of them it showed either minimal (e.g. 0.6% relative improvement on the Yeast data) or no improvement at all. The performance of the extended method measured in terms of two performance measures was lower on the Yeast data in comparison to the baseline classifiers and it was generally inferior or equal to them in more than half of all experiments (79 from 140).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Improvement with Rare Association Rules</head><p>Besides relatively low improvement demonstrated by the existing methods, they have the problem of using the standard support-confidence framework for AR mining, which normally extracts high support rules that often exist between large classes. So they ignore small classes as a potential source for improvement because minimum support filtering can remove not only noise but also rare classes. The greatest problem of Confidence in such a setup is that associations of small classes to large ones are normally ranked very high. In the case of a class ontology these ARs simply show hierarchical parent-child relations, i.e. that one label a is more specific than another b. The rule extracted would be a→b, which means that if label a appears, then label b should appear too. Such obvious relations can be derived from an extracted hierarchy, on the one hand, as in <ref type="bibr" target="#b2">[3]</ref> and are misleading for classification improvement, on the other. The reason being that applying such rules in the case of LRwAR <ref type="bibr" target="#b4">[5]</ref> for removing classes with a high support and setting them based on predictions of classes with a lower support, is prone to error. An example would be if class A appeared only 10 times, class B appeared 100 times and both appeared together 10 times, so a rule A→B would be extracted. LRwAR would then imply that B should be removed from the labelset and only in the post-processing step reinserted based on the prediction of A. Although in <ref type="bibr" target="#b3">[4]</ref> Conviction was used instead of Confidence, it is still closely related to Confidence and behaves similarly.</p><p>Another problem with the standard AR framework is choosing the thresholds for Support and Confidence which is done manually and can therefore be suboptimal. So, the first issue to be dealt with improving multi-label classification performance by constraints is the acquisition of high-quality rules. In the proposed approach we solve this problem by omitting the minimum support threshold and using a special interestingness measure which is well-suited for rare rules. Additionally we tune its threshold automatically depending on the range of values for extracted rules. The idea is to use rare ARs between classes that is from a small class to another small one and that classes belong to two or more ontologies describing different aspects of a dataset. In this case hierarchical relations between the classes of one ontology will not be taken into account. In such a setting, training data are annotated with categories of both ontologies. Then mined rules are used to improve class predictions for one of both ontologies. So, rare cross-ontology rules can be helpful in order to solve the described problems.</p><p>The second important issue is deciding when to apply a rule. Is the predicted label trusty enough to insert an additional label based on its presence? The worst case scenario would be that the rule is applied on the basis of a false positive inserting an additional false positive. Another undesirable outcome would be that the antecedent label is a true positive but applying a rule would create a false positive, i.e. the prediction of the classifier should not be overruled. The desirable decision is only to use a true positive to add another true positive, i.e. that a rule corrects the missclassification of a label. In <ref type="bibr" target="#b7">[8]</ref> the rules are applied to all labels in the rules, assuming that all antecedents were reliably predicted. Although the rules were used before to optimize the Ranking Loss, it was not clear if the antecedent's ranking was high enough to be predicted. In order to solve this problem, the control of the quality of classifier predictions is proposed to create a basis for application of a rule. The detailed discussion of the proposed approach is presented in the next two subsections.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">Selection of Rules</head><p>To extract pairwise ARs, we performed experiments with the interestingness measures well-suited for mining rare rules <ref type="bibr" target="#b13">[14]</ref>. Such measures should possess the important property of null-transaction invariance <ref type="bibr" target="#b0">[1]</ref>. Due to the lack of space we will focus here only on the Kulczynski measure (Kulc) which showed good results:</p><formula xml:id="formula_0">Kulc(A, B) = P AB 2 * ( 1 P A + 1 P B ).</formula><p>In order to select only the best rules, adaptive thresholding without a predefined value was applied as follows: After calculating Kulc of all rules, the values are sorted in descending order as a curve C. We assume that there will be a slope between a few high scored rules and the rest. In order to select these rules the curve C is smoothed into S and only the part with a relatively low variance is analyzed. Thus we need first to determine whether the variance of the curve S is high:</p><formula xml:id="formula_1">CV = MEAN (S) − VAR(S) MEAN (S) &gt; ρ m<label>(1)</label></formula><p>where MEAN (S) is the average value of the curve S and VAR(S) its variance.</p><p>If the condition of Eq. ( <ref type="formula" target="#formula_1">1</ref>) is true, we use only the values in the slope of the curve and calculate the median that defines a Threshold Value TV for the most interesting rules:</p><formula xml:id="formula_2">TV = C MEDIAN ({i|DIFF (Si)&gt;MEAN (DIFF (S))})</formula><p>where DIFF (S) is the difference between two neighbor values in the curve S.</p><p>Since the step size between two values is 1, DIFF (S) can also be seen as the derivative of S. Otherwise, i.e. if the variance is high, the average of the values not much lower than the mean of the entire curve is taken:</p><formula xml:id="formula_3">TV = MEAN (S j ) j∈{i|Si&gt;MEAN (S) * ρt}</formula><p>.</p><p>Defining the threshold in this manner, we select only those rules that have Kulc values above TV as good enough to be applied to prediction improvement.</p><p>ρ m and ρ t should be set so that the changes are significant and the only high valued rules are selected, respectively.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Application of Rules</head><p>As discussed above, applying a rule for insertion of a label without taking the corresponding classifier's judgment into account can lead to no improvement or even poorer prediction performance. A better way would be to use rankings produced by the classifier. An attempt was proposed in <ref type="bibr" target="#b7">[8]</ref> where the scores provided by the classifier for each label were used to optimize a parameter w varied from 0 to 1 for each pair of labels i and j. For a rule i→j, new rankings of label j were calculated for each sample x as: p j (x) = w * p j (x) + (1 − w) * p i (x), where p i (x) is the score assigned to a label i, analogously for j. by its respective BR classifier for that sample. These new rankings were used to minimize Ranking Loss by varying w during cross-validation on a validation set. However the label i was chosen in the test phase, only if its score was above a threshold t used to turn predicted rankings into classes (also called decision boundary later on). Thus the rule i→j could not be applied otherwise.</p><p>A drawback of this method is that it relies on individual parameter optimization for each rule through a cost-intensive calculation of the Ranking Loss. This is not viable for large datasets as the later work <ref type="bibr" target="#b5">[6]</ref> shows by using a fixed parameter value.</p><p>We propose a similar approach. The antecedent A of a rule A→B should be already positively predicted, i.e. should have a score greater than the threshold t, but an additional criterion should hold: V B V A &gt; 0.5, i.e. the score of the consequent should be at least 50% of the value of the antecedent in order to set the consequent.</p><p>As emphasized before, our approach was designed to work on classification problems with two different multi-label sets coming from two ontologies, but it can still be applied to the problems with only one class taxonomy in order to compare to other methods, as is shown in the next section.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Experiments</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1">Data</head><p>We used two multi-label real-world datasets: Reuters and Yeast. The first one was used with two class ontologies "Topics" and "Industries" for mining crossontology ARs as well as in a simplified version with only "Topics" labels in order to compare our method to the results of other improvement methods published elsewhere.</p><p>The two-ontology Reuters dataset was formed by preprocessing with stopword removal and stemming the original data provided by http://trec.nist. gov/data/reuters/reuters.html. We used the 5000 most frequent terms in the training set and applied tf-idf weighting as well as column-wise normalization performed separately on training and test data. In the original 800k samples only 300k contained at least one "Industries" label. From these we selected random 30k samples and split them into a training and a test set with the ratio of 2:1, i.e. 20k training samples and 10k test samples. In total there were 103 "Topics" labels and 364 "Industries" labels. We will denote this dataset as Reuters 10k below.</p><p>The simplified version (Reuters 5k with only "Topics" classes) consisted of 5000 training and 5000 test samples chosen randomly from the original 23k training set. The data preprocessing was performed as described above.</p><p>For the sake of comparison, the Yeast dataset was also taken from the MEKA package <ref type="bibr" target="#b10">[11]</ref>. It contains only 14 labels in one non-hierarchical label set and is therefore not very interesting for our experiments, but it is often used in the works on multi-label classification. From its 2417 samples, 1500 were selected randomly for training and the rest for test. We did not use cross-validation since certain aspects would be more difficult to analyze, for example, the graphs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2">Experiment Setting</head><p>As a baseline classifier we used LIBLINEAR <ref type="bibr" target="#b6">[7]</ref> in the BR setting, and on single class ontology datasets also ML-kNN as well as Classifier Chains (CC) based on LIBLINEAR. A crucial complication with LIBLINEAR is that the choice of the threshold t can be difficult. Normally, the value of 0.5 is recommended but in the work of <ref type="bibr" target="#b7">[8]</ref> a different value and individual for each dataset (0.45 for Yeast and 0.47 for Reuters) was chosen. We also used different values for each dataset and additionally compared the results to the results of an adaptive method for selecting t, which automatically adjusts it to be close to the dataset label cardinality <ref type="bibr" target="#b12">[13]</ref>. We will refer here to this method as Label Cardinality Approach (LCA). ML-kNN and CC were not used with two class ontologies because they do not scale well on large datasets, specially with LCS.</p><p>For LRwAR we also implemented a variation using rankings for all classes and only inserting new labels. We use the acronym OF (only fill) for this variation. The original method foresees deleting rankings of classes in the consequent of a rule as well.</p><p>Parameters of LRwAR and LCS were set as in their original works <ref type="bibr" target="#b4">[5,</ref><ref type="bibr" target="#b7">8]</ref>, whereas the Confidence threshold was used to obtain the best results for both methods. In particular, we changed it for Reuters 5k so that the h-loss performance measure was comparable to the value of the baseline classifier.</p><p>Parameters for adaptive thresholding were set to ρ m = 0.2 and ρ t = 0.75. These values were obtained by the manual experimentation on the Reuters dataset and are, in our opinion, general enough to be used for all datasets.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.3">Performance measures</head><p>We used the F-1 measure, which is the harmonic mean of Recall and Precision. It can be calculated in several ways depending on averaging <ref type="bibr" target="#b14">[15]</ref>. First, we used instance-based averaging, i.e. we calculated F-1 for every single instance and then took the mean value (denoted as IF1). Additionally we used label-based F-1 both in micro-averaged version mF1 and in macro-averaged one LF1=</p><formula xml:id="formula_4">1 Q i=1...Q 2 * tpi 2 * tpi+f ni+f pi .</formula><p>Here Q is the number of labels and tp i , f p i and f n i are, respectively, the number of true, false positives and false negatives for a label i. Micro-averaged mF1 is known to be dominated by the performance on large classes. Also Hamming Loss (h-loss) was used: HL = f p+f n Q * N , where N is the number of test samples.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.4">Results</head><p>Datasets with one class taxonomy: Yeast and Reuters 5k First we compared our approach to LRwAR,LCS, and LCA on the datasets used in other studies. In Table <ref type="table" target="#tab_0">1</ref> the results for Yeast and Reuters 5k are depicted.</p><p>On the Yeast data, IRAR, LCS, and LRwAR OF could improve the results of BR and ML-kNN classifiers in terms of mF1, IF1, and LF1. The improvement achieved by IRAR was the highest. H-loss for this dataset could not be increased by any of the compared methods. Among them LRwAR was the worst because its results were even worse as those of the baseline classifiers in terms of all performance measures. In contrast to the other methods, IRAR was also better than LCA for BR and comparable to it for ML-kNN. This shows that a powerful thresholding strategy can outperform many improvement methods based on label constraints. IRAR was the only improvement methods that could increase the CC results. This was the highest LF1 value by far on this dataset. This is consistent with the important fact that IRAR could increase the LF1 value significantly more than the other methods in almost all configurations. The only exception was for Reuters 5k and BR where it improved second best. This can be explained by the better improvement of the classification performance on small classes. Indeeed, as Figure <ref type="figure" target="#fig_0">1a</ref> shows, the number of true positives on small and middle-size classes obtained by IRAR was higher than that of LCA (Figure <ref type="figure" target="#fig_0">1b</ref>). This difference is even more pronounced if we compare F-1 values for each class obtained by all improvement methods and presented in Figure <ref type="figure" target="#fig_1">2a</ref>. One can see, for example, that IRAR achieves a significant improvement in F-1 for the last class where the other methods show no improvement at all or that it has much more improvement on the classes 5-10.</p><p>Analyzing the curves of mF1 and LF1 in dependence on the threshold t one can see a trade-off between them (Figure <ref type="figure" target="#fig_1">2b</ref>). IRAR is able to achieve both high LF1 and mF1 values near their crossing point.</p><p>On the Reuters 5k dataset, IRAR had again the highest improvement against the baseline classifier as compared to the other improvement methods in terms of all performance measures, except for h-loss. The largest performance difference was again in LF1. IRAR performance was comparable to that of LCA. LCS and LRwAR achieved a very small improvement against the baseline classifier and were worse than LCA in terms of all performance measures, except for h-loss.</p><p>Here we can see that CC had was the second best classifier, but no improvement could beat LCA method. Again, the exception remains IRAR with LF1, having a 18% value increase over the baseline performance and 3% over LCA.</p><p>CC did not outperform BR in the experiments, although CC does consider the connections between the labels in a certain way. A solution would be to use Ensembles of CC (ECC) <ref type="bibr" target="#b11">[12]</ref>, since the order of the labels can be taken into account. However for ECC, the issue of larger label sets will be even much severe, since the label order must be permutated when creating a new CC to exhaust all alternatives at best.   Dataset with two class ontologies: Reuters 10k Table <ref type="table" target="#tab_1">2</ref> depicts the results of classification improvement for the Reuters 10k dataset, first classified separately in "Topics" and "Industries" and then with improved "Industries" predictions, by using cross-ontological ARs. In general, the classification performance for "Topics" was higher than for "Industries" classes. The results of the improvement methods LCS and IRAR for this class ontology were better than those of the baseline classifier, except that h-loss of IRAR was lower. At the same time, LRwAR showed negative improvement and LRwAR OF only improvement at the fourth place after the decimal point. In contrast, IRAR was able to achieve the overall best LF1. Its results were also somewhat similar to the results of LCA. It is interesting to note that LCS outperformed LCA in terms of mF1 and IRAR in terms of LF1. So, we can conclude that LCS is more effective for classes with large support and IRAR for those with small support. This will be due to the use of confidence to extract the rules.</p><p>The results for Reuters 10k "Industries" are similar to those obtained for "Topics". Here LRwAR had even more negative improvement in terms of all performance measures and LRwAR OF showed again only marginal improvement. LCS was equal or better than the baseline and achieved again the highest mF1 value. This time both LCA and IRAR were worse than the baseline in terms of h-loss and mF1, but improved IF1 and LF1. However IRAR was better than LCA in three out of four performance measures and had again the best LF1.</p><p>Using cross-ontology ARs for the improvement of "Industries" predictions revealed an interesting fact: the results of LCS and both LRwAR variants worsened in comparison with those shown in the previous experiment while IRAR could improve its h-loss and mF1 values. Here, LCS uses the thresholds of different classifiers trained with different labelsets that may obstruct its performance. Also the low occurrence of labels in the labelsets may lead to poor results of the Confidence-based methods. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">Conclusion</head><p>In this paper we proposed a novel method of classification improvement in multilabel classification IRAR. It uses cross-ontology association rules and focuses on the improvement of predictions for small classes. Additionally, we compared it with state-of-the-art methods developed to correct predicted rankings by using constraints on labels in a post-processing step. One of the methods, LRwAR, showed negative improvement in most of the experiments and its variation only marginal improvement. LCS scored better in terms of improvement, but a better thresholding strategy such as LCA often achieved even more improvement. IRAR could outperform LCA in three out of four performance measures on the Yeast and Reuters 10k datasets. More importantly, it boosted the LF1 value significantly and showed the best LF1 result in three of five experiments. This means that IRAR is well suited for improving performance on small classes. This method is also able to achieve the trade-off between LF1 and mF1, i.e. it was able to achieve a high LF1 at a relatively low number of false positives. This points to the fact that the method can be used effectively with datasets exhibiting highly skewed label distributions as, for example, in the case of class ontologies.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>IRARFig. 1 .</head><label>1</label><figDesc>Fig. 1. Distributions of true positives on Yeast data.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Fig. 2 .</head><label>2</label><figDesc>Fig. 2. Improvement comparison on Yeast data.</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1 .</head><label>1</label><figDesc>LCS, LRwAR, and IRAR applied to Yeast and Reuters 5k, OF=Only Filling, t = threshold, bold values mark the best values per dataset and column.</figDesc><table><row><cell></cell><cell>BR</cell><cell></cell><cell></cell><cell cols="2">ML-kNN</cell><cell></cell><cell>CC</cell><cell></cell><cell></cell></row><row><cell>Metrics</cell><cell>h-loss mF1</cell><cell>IF1</cell><cell>LF1</cell><cell>h-loss mF1</cell><cell>IF1</cell><cell>LF1</cell><cell>h-loss mF1</cell><cell>IF1</cell><cell>LF1</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell>Yeast</cell><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell></row><row><cell>LCA</cell><cell cols="9">0.2121 0.6498 0.6362 0.4088 0.2066 0.6583 0.6467 0.4042 0.2148 0.6455 0.6326 0.4053</cell></row><row><cell></cell><cell cols="2">t=0.45</cell><cell></cell><cell cols="2">t=0.5</cell><cell></cell><cell cols="2">t=0.45</cell><cell></cell></row><row><cell>baseline</cell><cell cols="9">0.2043 0.6477 0.6288 0.3915 0.1990 0.6221 0.5978 0.3496 0.2097 0.6370 0.6184 0.3854</cell></row><row><cell>LCS Cnf=0.6</cell><cell cols="9">0.2071 0.6516 0.6326 0.3983 0.1996 0.6263 0.6016 0.3596 0.2141 0.6313 0.6144 0.3540</cell></row><row><cell>LRwAR Cnf=0.6</cell><cell cols="9">0.2071 0.6328 0.6160 0.3479 0.2047 0.6034 0.5838 0.2982 0.2121 0.6223 0.6057 0.3401</cell></row><row><cell cols="10">LRwAR OF Cnf=0.6 0.2047 0.6480 0.6293 0.3935 0.1990 0.6221 0.5978 0.3496 0.2105 0.6367 0.6183 0.3857</cell></row><row><cell>IRAR</cell><cell cols="9">0.2269 0.6544 0.6400 0.4453 0.2169 0.6560 0.6446 0.4101 0.3553 0.6031 0.5992 0.4760</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell>Reuters 5k</cell><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell></row><row><cell>LCA</cell><cell cols="9">0.0131 0.7893 0.7955 0.4177 0.0164 0.7361 0.7446 0.4306 0.0148 0.7608 0.7716 0.3910</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell>t=0.3</cell><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell></row><row><cell>baseline</cell><cell cols="9">0.0122 0.7849 0.7817 0.3690 0.0164 0.7339 0.7376 0.4303 0.0133 0.7567 0.7492 0.3302</cell></row><row><cell cols="10">LCS n=6,Cnf=0.8 0.0122 0.7856 0.7823 0.3696 0.0165 0.7335 0.7377 0.4311 0.0140 0.7382 0.7305 0.3246</cell></row><row><cell cols="10">LCS n=6, Cnf=.85 0.0122 0.7856 0.7823 0.3696 0.0165 0.7335 0.7377 0.4311 0.0140 0.7383 0.7307 0.3249</cell></row><row><cell>LRwAR Cnf=0.8</cell><cell cols="9">0.0138 0.7465 0.7442 0.3576 0.0177 0.7002 0.7015 0.4191 0.0148 0.7170 0.7119 0.3194</cell></row><row><cell cols="10">LRwAR Cnf=0.85 0.0130 0.7667 0.7633 0.3614 0.0169 0.7189 0.7204 0.4231 0.0140 0.7377 0.7302 0.3231</cell></row><row><cell cols="10">LRwAR OF Cnf=0.8 0.0122 0.7851 0.7819 0.3690 0.0164 0.7340 0.7378 0.4303 0.0132 0.7584 0.7508 0.3313</cell></row><row><cell cols="10">LRwAR OF Cnf=0.85 0.0122 0.7849 0.7817 0.3690 0.0164 0.7339 0.7376 0.4303 0.0132 0.7584 0.7508 0.3313</cell></row><row><cell>IRAR</cell><cell cols="9">0.0125 0.7900 0.7895 0.3958 0.0187 0.7174 0.7347 0.4433 0.0164 0.7452 0.7490 0.4001</cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 2 .</head><label>2</label><figDesc>LCS, LRwAR, and IRAR applied to Reuters BR's predictions for "Topics" and "Industries" 10k, OF=Only Filling, t = threshold, bold values mark the best values per dataset and column. Topics" t=0.45 Reuters 10k "Industries" t=0.3 LCA 0.0123 0.8257 0.8335 0.4237 0.0070 0.6462 0.6466 0.2884 baseline 0.0116 0.8258 0.8282 0.3938 0.0061 0.6589 0.6060 0.2772 LCS k=5,n=6,Cnf=0.7 0.0116 0.8264 0.8287 0.3942 0.0061 0.6605 0.6077 0.2778 LCS k=5,n=6,Cnf=0.85 0.0116 0.8264 0.8287 0.3942 0.0061 0.6603 0.6076 0.2776 LRwAR Cnf=0.7 0.0134 0.7912 0.7890 0.3819 0.0071 0.5491 0.4649 0.2672 LRwAR Cnf=0.85 0.0122 0.8143 0.8145 0.3871 0.0067 0.5936 0.5138 0.2698 LRwAR OF Cnf=0.7 0.0116 0.8259 0.8284 0.3940 0.0061 0.6592 0.6068 0.2773 LRwAR OF Cnf=0.85 0.0116 0.8260 0.8285 0.3940 0.0061 0.6592 0.6067 0.</figDesc><table><row><cell>Metrics</cell><cell>h-loss mF1</cell><cell>IF1</cell><cell>LF1</cell><cell>h-loss mF1</cell><cell>IF1</cell><cell>LF1</cell></row><row><cell></cell><cell cols="6">Reuters 10k "2773</cell></row><row><cell>IRAR</cell><cell cols="6">0.0132 0.8187 0.8312 0.4298 0.0067 0.6539 0.6120 0.2918</cell></row><row><cell cols="3">Reuters 10k "Topics"→"Industries", t=0.3</cell><cell></cell><cell></cell><cell></cell><cell></cell></row><row><cell>LCS Cnf=0.7,</cell><cell cols="3">0.0061 0.6589 0.6060 0.2772</cell><cell></cell><cell></cell><cell></cell></row><row><cell cols="4">LCS k=5,n=6,Cnf=0.85, 0.0061 0.6589 0.6060 0.2772</cell><cell></cell><cell></cell><cell></cell></row><row><cell>LRwAR Cnf=0.7</cell><cell cols="3">0.0087 0.3482 0.2880 0.2293</cell><cell></cell><cell></cell><cell></cell></row><row><cell cols="4">LRwAR OF Cnf=0.7 0.0061 0.6585 0.6056 0.2771</cell><cell></cell><cell></cell><cell></cell></row><row><cell>IRAR</cell><cell cols="3">0.0062 0.6590 0.6092 0.2825</cell><cell></cell><cell></cell><cell></cell></row></table></figure>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Evaluation of hierarchical interestingness measures for mining pairwise generalized association rules</title>
		<author>
			<persName><forename type="first">F</forename><surname>Benites</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Sapozhnikova</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">IEEE Trans. Knowl. Data Eng</title>
		<imprint>
			<biblScope unit="volume">26</biblScope>
			<biblScope unit="issue">12</biblScope>
			<biblScope unit="page" from="3012" to="3025" />
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">Mining rare associations between biological ontologies</title>
		<author>
			<persName><forename type="first">F</forename><surname>Benites</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Simon</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Sapozhnikova</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">PLoS ONE</title>
		<imprint>
			<biblScope unit="volume">9</biblScope>
			<biblScope unit="page">e84475</biblScope>
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Multi-label classification and extracting predicted class hierarchies</title>
		<author>
			<persName><forename type="first">F</forename><surname>Brucker</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Benites</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><forename type="middle">P</forename><surname>Sapozhnikova</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Pattern Recognition</title>
		<imprint>
			<biblScope unit="volume">44</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="724" to="738" />
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">LI-MLC: A label inference methodology for addressing high dimensionality in the label space for multilabel classification</title>
		<author>
			<persName><forename type="first">F</forename><surname>Charte</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Rivera</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Del Jesus</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Herrera</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">IEEE Trans. Neural Netw. Learn. Syst</title>
		<imprint>
			<biblScope unit="volume">25</biblScope>
			<biblScope unit="issue">10</biblScope>
			<biblScope unit="page" from="1842" to="1854" />
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">Improving multi-label classifiers via label reduction with association rules</title>
		<author>
			<persName><forename type="first">F</forename><surname>Charte</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Rivera</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Del Jesus</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Herrera</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Hybrid Artificial Intelligent Systems</title>
				<imprint>
			<date type="published" when="2012">2012</date>
			<biblScope unit="volume">7209</biblScope>
			<biblScope unit="page" from="188" to="199" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">Improving multi-label classification performance by label constraints</title>
		<author>
			<persName><forename type="first">B</forename><surname>Chen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Hong</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Duan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Hu</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">IJCNN</title>
		<imprint>
			<biblScope unit="page" from="1" to="5" />
			<date type="published" when="2013-08">2013. Aug 2013</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<analytic>
		<title level="a" type="main">Liblinear: A library for large linear classification</title>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">E</forename><surname>Fan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><forename type="middle">W</forename><surname>Chang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><forename type="middle">J</forename><surname>Hsieh</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><forename type="middle">R</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><forename type="middle">J</forename><surname>Lin</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">JMLR</title>
		<imprint>
			<biblScope unit="volume">9</biblScope>
			<biblScope unit="page" from="1871" to="1874" />
			<date type="published" when="2008">2008</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<analytic>
		<title level="a" type="main">Combining binary-svm and pairwise label constraints for multi-label classification</title>
		<author>
			<persName><forename type="first">W</forename><surname>Gu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Chen</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Hu</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Systems Man and Cybernetics (SMC)</title>
				<imprint>
			<date type="published" when="2010">2010. 2010</date>
			<biblScope unit="page" from="4176" to="4181" />
		</imprint>
	</monogr>
	<note>IEEE International Conference on</note>
</biblStruct>

<biblStruct xml:id="b8">
	<analytic>
		<title level="a" type="main">Interestingness measures and strategies for mining multi-ontology multi-level association rules from gene ontology annotations for the discovery of new go relationships</title>
		<author>
			<persName><forename type="first">P</forename><surname>Manda</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><forename type="middle">M</forename><surname>Mccarthy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">M</forename><surname>Bridges</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">J. of Biomedical Informatics</title>
		<imprint>
			<biblScope unit="volume">46</biblScope>
			<biblScope unit="issue">5</biblScope>
			<biblScope unit="page" from="849" to="856" />
			<date type="published" when="2013">2013</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Multi-label classification with label constraints</title>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">H</forename><surname>Park</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Fürnkranz</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">ECML PKDD 2008 Workshop on Preference Learning</title>
				<imprint>
			<date type="published" when="2008">2008</date>
			<biblScope unit="page" from="157" to="171" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<monogr>
		<author>
			<persName><forename type="first">J</forename><surname>Read</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Reutemann</surname></persName>
		</author>
		<ptr target="http://meka.sourceforge.net/" />
		<title level="m">Meka multi-label dataset repository</title>
				<imprint>
			<date type="published" when="2015-05-20">May 20 2015</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Classifier chains for multi-label classification</title>
		<author>
			<persName><forename type="first">J</forename><surname>Read</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Pfahringer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Holmes</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Frank</surname></persName>
		</author>
		<idno type="DOI">10.1007/978-3-642-04174-7_17</idno>
		<ptr target="http://dx.doi.org/10.1007/978-3-642-04174-7_17" />
	</analytic>
	<monogr>
		<title level="m">European Conference on Machine Learning and Knowledge Discovery in Databases: Part II</title>
				<meeting><address><addrLine>Berlin, Heidelberg</addrLine></address></meeting>
		<imprint>
			<publisher>Springer-Verlag</publisher>
			<date type="published" when="2009">2009</date>
			<biblScope unit="page" from="254" to="269" />
		</imprint>
	</monogr>
	<note>ECML PKDD &apos;09</note>
</biblStruct>

<biblStruct xml:id="b12">
	<analytic>
		<title level="a" type="main">Classifier chains for multi-label classification</title>
		<author>
			<persName><forename type="first">J</forename><surname>Read</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Pfahringer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Holmes</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Frank</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Machine learning</title>
		<imprint>
			<biblScope unit="volume">85</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="333" to="359" />
			<date type="published" when="2011">2011</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b13">
	<analytic>
		<title level="a" type="main">Selecting a right interestingness measure for rare association rules</title>
		<author>
			<persName><forename type="first">A</forename><surname>Surana</surname></persName>
		</author>
		<author>
			<persName><forename type="first">U</forename><surname>Kiran</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><forename type="middle">K</forename><surname>Reddy</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">16th Int. Conf. on Management of Data (COMAD)</title>
				<imprint>
			<date type="published" when="2010">2010</date>
			<biblScope unit="page" from="115" to="124" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b14">
	<analytic>
		<title level="a" type="main">Mining multi-label data</title>
		<author>
			<persName><forename type="first">G</forename><surname>Tsoumakas</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Katakis</surname></persName>
		</author>
		<author>
			<persName><forename type="first">I</forename><surname>Vlahavas</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Data Mining and Knowledge Discovery Handbook</title>
				<imprint>
			<date type="published" when="2010">2010</date>
			<biblScope unit="page" from="667" to="685" />
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
