<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Information Technologies and Security, December</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Bulgarian Defense Institute</institution>
          ,
          <addr-line>"Professor Tzvetan Lazarov" Blvd. 2 1592 Sofia</addr-line>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Information Recording of NAS of Ukraine</institution>
          ,
          <addr-line>Shpak Str. 2 03113 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>19</volume>
      <issue>2024</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>An approach to identifying information objects (IOs), data about which is received by the monitoring system from independently operating sources, is presented. It considers a situation where data about the same physical object can be entered multiple times into an information resource, as about different objects. At the same time, the values of such IOs features do not completely coincide, since the data sources introduce some operation errors. The proposed approach for object identification is based on a new proximity (similarity) measure of information objects, which takes into account the existing probabilistic uncertainty regarding the values of quantitative features and the uncertainty of the possibility type for qualitative features.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;identification</kwd>
        <kwd>information object</kwd>
        <kwd>monitoring system</kwd>
        <kwd>proximity (distance) measure</kwd>
        <kwd>quantitative features</kwd>
        <kwd>qualitative features</kwd>
        <kwd>probability distribution law</kwd>
        <kwd>fuzzy set 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The majority of information systems have a functionally, and sometimes organizationally
dedicated monitoring subsystem, the aim of which is to extract (obtain) data about the external
environment and the state of the system as a whole. This work focuses on one of the problematic
tasks of forming an information resource that is populated with data from monitoring the external
environment, shaped by the actions of dynamic objects in the surrounding space. As the number of
monitored objects increases, along with expanding the means of monitoring, or when multiple
monitoring systems are integrated into a higher-level system, there is an increased probability that
data about the same object can independently be entered into the common information resource of
the monitoring system. This situation is typical for cases when the monitoring system (or
subsystem, if the monitoring system is hierarchical) simultaneously analyzes objects in the
surrounding space in overlapping spatial areas. Figure 1 schematically shows several overlapping
areas monitored by different data sources (DS), and, accordingly, a few sources can observe the
same physical (real) objects. The set of features of such objects, determined by the data source, will
be called an information object (IO). Essentially, the IO, formed by a suitable data source, is an
information representation of a real object in the system’s information resource in the form of a
finite set of features and their values.</p>
      <p>In the case described above, there is a need to solve the problem of identifying IOs, that is,
attributing them to a single physical object, with the subsequent unification (aggregation) of the
characteristics of such (identified) IOs. In the simplest case, the values of IO features that refer to
the same real object should completely coincide, even if they are determined by different data
sources. Such a coincidence would make it quite easy to classify a certain number of IOs as those
describing the same physical object, provided that in the area of the monitoring system operation,
there are no different objects with completely identical observational characteristics. Regarding the
latter condition, we would additionally note that if the monitoring system allows the formation of
completely identical IOs without additional (marking) data that refer to different physical objects,
this indicates its insensitivity to certain differences in the external environment, which can be seen
as a flaw.</p>
      <p>
        In our case, given that among the key object features in the surrounding space are the static or
time-varying coordinates of their location, we can reasonably assume that the condition above is
satisfied. However, even if the feature values for different IOs referring to the same real object are
equal, there would be a need to identify such IOs due to the fact that not all sources are able to
determine the full list of features. Therefore, even in such a simple case, to solve the problem of
identifying the IO, it would be necessary to introduce and analyze a certain IO proximity measure,
such as Rao coefficients [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Additionally, the situation is further complicated by the fact that the
monitoring system tools determine any features of objects with errors, which makes the chance of
an exact match of even some feature values random and unlikely. This prompts a transition from
searching for an exact equality of IO features to analyzing the proximity of IOs across the full set of
features available for observation, taking into account the errors in their determination. Therefore,
solving the problem of identifying IOs requires solving the problem of formally defining the
proximity measure between the feature values and between IOs as a whole.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        The problem of determining which observations or descriptions correspond to the same object
(object instance) exists in various fields. These can be image recognition and analysis, natural
language generation and processing, text processing, integration of information resources, etc. For
example, to track a moving object in computer vision systems, it is necessary to identify whether
two shapes in different frames of a video stream are actually the same object. Creating a distributed
information system with a common information space involves merging separate databases. As a
result, it is necessary to determine which records belong to the same entity and solve the
integration task when incomplete matching of their attributes. When preparing reference lists in
articles, it is necessary to find which citations refer to the same papers to avoid duplication. In
natural language processing, a key task is to determine which phrases (word combinations) refer to
the same entity. An object identification for databases data merging/cleansing, record linking and
duplicate removal was first formulated as a separate problem by Newcombe et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and solved by
      </p>
      <sec id="sec-2-1">
        <title>Fellegi and Sunter [3], whose method became the basis for further developments.</title>
        <p>
          There are currently numerous developments in this area, including Wang and Ji [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], Nagarajan
and Grauman [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], Singla and Domingos [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], Cohen and Richman [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], Sarawagi and Bhamidipaty
[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], Pasula et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], etc. Most existing approaches are the development and improvement of the
original Fellegi and Sunter model, in which object identification is considered as a classification
problem. That is, it defines a vector of similarity scores between the attributes of two observations,
based on which the classification into "Match" or "Not Match" is performed. Each candidate pair is
assessed separately, and a matching decision is formed. Then a transitive closure is constructed to
eliminate inconsistencies. At the same time, the development of new methods is ongoing in two
directions: improving measures and metrics for assessing the proximity of research objects and
improving methods for group processing of multiple comparisons.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Problem formulation</title>
      <p>The set of information objects that will be used by an information system is first defined during its
design phase. Later, over the life of the system, this set is supplemented and edited in accordance
with users’ information needs. IOs can describe:
for 


the set of features of information objects set  ,  = 1 …</p>
      <p>– association relationship (semantic relations).




</p>
      <sec id="sec-3-1">
        <title>Single entities (material objects, persons) Abstract entities (concepts)</title>
      </sec>
      <sec id="sec-3-2">
        <title>Group entities (homogeneous or heterogeneous)</title>
      </sec>
      <sec id="sec-3-3">
        <title>Static composite entities (situation description)</title>
      </sec>
      <sec id="sec-3-4">
        <title>Dynamic composite entities (processes). An information object can be formally specified by a tuple where m – IO unique identificator</title>
        <p>= 〈 ,  ,  ,  〉,
 = 〈 ,  ,  〉,
 =   = 1 …  . – a set of IO features (attributes)
 = { | = 1 …  }. – a set of constraints on the object attribute values
 : →  . – a mapping to set constraints for each attribute.</p>
        <p>In general, information objects are interconnected and interdependent. Let us define the set of
relations between IOs as follows:
where 
– inheritance relationship ("class-subclass")  (
, 
), where 
is an upper class
– aggregation relationship ("included in") 
, 
, where 
features are included in
(1)
(2)</p>
        <p>
          The total set of possible IO features can be divided into features of a quantitative and qualitative
nature. The values of quantitative features are determined using certain measuring instruments
and, accordingly, are characterized by uncertainty of the type "probability", since any means of
measurement has limited, albeit defined, accuracy. The values of qualitative features are
determined by the active participation of a person, and therefore contain a subjective component,
which today is usually described by uncertainty of the type "possibility" [
          <xref ref-type="bibr" rid="ref10 ref11 ref15">10, 11, 15</xref>
          ]. It is obvious
that in this case, the proximity (or distance) measure between the IOs described by a set of features
should be a combined quantitative-qualitative one. Few such proximity measures are known. In
particular, these include the Voronin approximation proximity measure [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], the Mirkin similarity
measure [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], and the most “physically transparent” Zhuravlev measure [
          <xref ref-type="bibr" rid="ref1 ref14">1, 14</xref>
          ], which for two IOs i
and j is determined as follows
(3)

=
        </p>
        <p>,
where  – an index of an object feature ( = 1,  ), L – total number of features;


=
=
1,  
0,  
− 
ℎ 
1,</p>
        <p>ℎ 
0,   ℎ 
≤  ,
;
 
 
(for quantitative features)</p>
        <p>(for qualitative features )
 - quantitative proximity threshold for the  -th feature.</p>
        <p>As can be seen from expression (3), Zhuravlev's measure for quantitative features allows the
possibility of some slight difference in their values, within which it is assumed that the features
still coincide. That is, a threshold analysis of the proximity of such feature values is used. It should
be noted that determining the threshold value 
when solving a specific problem is up to the
researcher. Such a possibility is not provided for qualitative features, and only complete
coincidence/difference of their values is allowed.</p>
        <p>Let us consider the acceptability of the approach to accounting for the possible difference in the
quantitative feature values 
− 
=</p>
        <p>through defining some admissible value of it -  . It is
known that measurement errors of various quantities are distributed according to a certain law of
probability distribution. This distribution is characterized by the mathematical expectation of the
error (the average error value, which is equal to zero in the absence of a systematic component),
the standard deviation from its mathematical expectation (mean square error - MSE), and other
higher-order moments.</p>
        <p>It is generally accepted that measurement errors are most often distributed according to a
normal law. Let us assume this statement is true for all our cases. Then the measurement error
distribution for quantitative features is completely determined by the first order moment
(mathematical expectation) and the second order moment - the dispersion, or the standard
deviation of the random variable. It is quite obvious that when the linear distance 
between the
measured feature values decreases, the probability that the obtained measured values actually refer
to the same true value will increase nonlinearly, in accordance with the distribution laws of their
measurement errors with two different means (which are generally characterized by different
MSE). In addition, if the measured feature values coincide, but the values were obtained with an
error, then such a coincidence cannot be guaranteed to mean a coincidence of the true values.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Therefore, using the constant  is a fairly rough approximation to reality.</title>
      </sec>
      <sec id="sec-3-6">
        <title>Therefore, this work aims to construct a proximity (similarity) measure to compare information objects, which takes into account the possibility of errors for both types of IOs features – quantitative and qualitative. 86</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Determining the proximity measure for quantitative features</title>
      <p>This paper proposes an alternative method for calculating the degree of proximity between the
quantitative feature values, which takes into account the probabilistic nature of the process of their
definition by different sources. The currently known proximity measures between two measured
(quantitative) feature values require calculating the probability that the true value of the feature for
both measurements is actually the same value. Calculating such a probability requires full
knowledge of the probability distribution laws for measured values (not just the measurement
errors) and, therefore, the true value of the measured quantity. In our problem, this value is
unknown. Taking into account the above, it is proposed to formalize the quantitative feature values
by the normal distribution law of their determination errors, where the feature value obtained from
the source is considered as the mathematical expectation. Data on the standard measurement error
is expected to be obtained from the data source or determined based on its characteristics as a
measurement means. Then the coincidence of the feature values obtained from two sources can be
considered dependent on the probability of finding the true value in the intersection area of the
two distribution laws, which can be calculated based on the Laplace function, the probability
multiplication theorem for independent events and the well-known "three sigmas" rule. The
expression for calculating the probability of finding a random variable x in the interval (с, d) with
its normal distribution has the following form
(4)
−
− Ф
 − 

,
where
Ф() – Laplace function
m – mathematical expectation of a random variable.</p>
      <sec id="sec-4-1">
        <title>Given the independence of the two measurements, the probability that the measured quantity is</title>
        <p>actually within the range of values ( ≤  ≤  ): 
=  ( , ) ∙  ( , ), where  ( , ) and  ( , ) - the
probabilities that the feature values for each measurement are within the interval (с, d).</p>
        <p>Consider the example shown in Figure 2. Let the value of the attribute x be measured by two
different sources. The obtained measurement results are: 
= 12 and 
= 18 units. In this case,
the MSE for measurement errors for each source are: 
= 3, 
= 2, and 
= 18 − 12 = 6.</p>
      </sec>
      <sec id="sec-4-2">
        <title>The boundaries of the overlapping regions of the probable value's areas (which determine the</title>
        <p>specified probability) for two variables - 
(taking into account the "three sigmas" rule) are
determined as the difference between the smaller value from the pair 
+ 3
and the larger value from the pair 
For the case shown in Figure 2: 
− 3
+ 3
and 
= 21; 
− 3 , where 
+ 3</p>
        <p>Next, we calculate the probability that the true value of each measurement is in the range 
from two data sources.</p>
      </sec>
      <sec id="sec-4-3">
        <title>We will perform similar calculations for other measured values of the same feature</title>
        <p>= 14 and
= 17 units (other parameters are the same as in the previous case), 
= 17 − 14 = 3.</p>
        <p>For this case: 
+ 3
= 23; 
+ 3
= 23; 
− 3
= 5; 
− 3
= 11 and,
respectively,</p>
        <p>= (11,23). Next we calculate</p>
        <p>(11 ≤  ≤ 23) = Ф
(11 ≤  ≤ 23) = Ф
2</p>
        <p>3
23 − 14</p>
        <p>11 − 14
23 − 17</p>
        <p>11 − 17
− Ф
− Ф
2
3</p>
        <p>= 0,84,
= 0,9973.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Therefore</title>
        <p>= 0,8377.</p>
        <p>Let's perform another calculation to analyze the change nature in the proposed proximity
measure and set 
= 15 and 
= 17 units (
= 2). Then: 
+ 3</p>
      </sec>
      <sec id="sec-4-5">
        <title>We also note that if the distribution laws and the measured values completely coincide, the</title>
        <p>probability calculated by the expression (1):</p>
        <p>≈ 1. If they do not intersect within 3, the value
will be equal to zero.</p>
        <p>Comparing the calculation results, it can be stated that as the difference between the two
feature value measurements</p>
        <p>obtained from different data sources decreases, the value of the
proximity measure increases. Additionally, the measure value changes nonlinearly in relation to
the linear change of</p>
        <p>
          in accordance with the distribution laws of measurement errors of the
feature value. Testing the proposed measure for compliance with known conditions for its
acceptability and validity (non-negativity, symmetry, maximum similarity of an object to itself, and
the “triangle inequality” [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]) shows the possibility of non-fulfilment of the last condition while
meeting the first three conditions. At the same time, the latter condition is considered additional
and optional [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-6">
        <title>The obvious advantage of using probability to calculate a proximity measure for quantitative</title>
        <p>features is that the probability (and therefore the measure) changes from 0 to 1, i.e., it is
immediately normalized. That is, calculating proximity values for different features does not
require their transformation when determining the common proximity for all features.</p>
      </sec>
      <sec id="sec-4-7">
        <title>Thus, the proposed proximity measure for quantitative features can be considered acceptable,</title>
        <p>especially since it is based on a probability analysis that has a specific physical meaning. Also, this
proximity measure is simply transformed into a distance measure by its inversion</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Determining the proximity measure for qualitative features</title>
      <p>Let us consider the problem of calculating the possibility that two specified values of a certain
qualitative feature are actually the same value. At the same time, it is important to remember that a
qualitative characteristic can be expressed in numbers and still retain its qualitative character, since
this feature character is determined not by the form of its expression (reflection), but by the
method of its acquisition.</p>
      <p>To determine the proximity measure between the values of qualitative features, it is proposed to
use their formalization in the form of fuzzy sets by constructing a triangular membership function
for each obtained feature value on a clear set of its possible values (an example is shown in Figure
3). The number of feature values in the support sets is determined by the possibility of other
feature values in reality (which can be considered a certain analogue of measurement error for
quantitative attributes). Then the distance between the qualitative feature values can be defined as
the maximum value of the set, which is the intersection of two fuzzy sets (formalized feature
values). So the expression for the proximity measure is
(5)
(6)
the form of fuzzy sets
where
 and</p>
      <p>– clear values of the feature x
() ,</p>
      <p>() – membership functions of fuzzy sets constructed for both values of
the qualitative feature.</p>
      <p>= 18. The region of clear feature values, which are covered by
non-zero values of the membership function, is specified by the maximum possible errors in the
feature determination. As a result, the proximity measure is between the obtained values 
=
0,67. It is obvious that if the clear feature values approach each other, the value of the proximity
measure approaches 1. Otherwise, it approaches zero.</p>
      <sec id="sec-5-1">
        <title>To obtain a distance measure, by analogy with quantitative features, it is also necessary to invert the obtained value, since as the obtained feature values approach each other, the value  will increase: the proximity measure is</title>
        <p>= 1 −</p>
      </sec>
      <sec id="sec-5-2">
        <title>The problem is solved similarly in relation to qualitative IO features, given by linguistic concepts on an ordinal scale. Then the fuzzy set is formed based on the term-set. In this case, to form fuzzy sets, it is necessary to sort the feature values by increasing (strengthening) of the object property that it characterizes.</title>
      </sec>
      <sec id="sec-5-3">
        <title>If the feature is nominal, the membership function will be characterized by one extreme value</title>
        <p>and some constant value  for all other members of the set, which will characterize the possibility
of false determination of the feature. Therefore, if the IO feature values obtained from two sources
do not coincide, their proximity will be determined by this value , regardless of the feature values</p>
      </sec>
      <sec id="sec-5-4">
        <title>The proposed proximity measure for qualitative features of fuzzy sets meets all four conditions</title>
        <p>Note, if, instead of the error distribution laws, we use their triangular approximation for
quantitative characteristics, the triangle inequality condition will also be met.
6. Determining the proximity measure (metric) for IOs on the set of
themselves.
for the measures’ validity.</p>
        <p>their features


(7)
(8)
To determine the metric (function) of IOs similarity by all features, we can use one of the known
additive functions with normalization by the number of values of quantitative and qualitative
features and, if necessary, different weight values for each feature, or subsets of quantitative and
qualitative features as a whole. It is desirable that the sum of the weight coefficients be equal to 1.
Considering the different natures of uncertainty for quantitative and qualitative features, the most
acceptable expression for determining the distance measure between IOs on the set of their
features is
= 
 – the weight coefficients for quantitative features.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7. Analysis of the IOs proximity metric over time to improve the quality of identification</title>
      <p>The final stage of solving the identification problem may be the analysis of the IOs’ behavior
(actions) over time from the point of view of possible changes in their feature values. This analysis
requires setting a criterion by which IOs are considered to be identified as a single object, given the
variable distance (proximity) between them over time. Such a criterion will be determined by the
characteristics of the chosen distance metric and the specific application problem. For example, if
the same physical object is observed by two data sources over several cycles of information update,
increasing the probability of correct object identification can be achieved by analyzing the linear
trend ( ) of the change in the distance metric (proximity). For this purpose, the well-known least
squares method (LSM) can be applied. Let us denote the distance metric value between two IOs at
numbers
natural numbers
we get
and the expression for the sum of the squares of the first N terms of the arithmetic progression of
 =

 =</p>
      <p>( + 1)
 ( + 1)(2 + 1)</p>
      <p>,
2
6
(9)
(10)
(11)
(12)
(13)
time i as  . Then, using the expression obtained based on LSM to calculate the trend line slope
and substituting successive time stamps from the interval  = 1,  into it, we obtain

= 
 
−



 −</p>
      <p>Using expressions for the sum of the first N terms of the arithmetic progression of natural
 = 
 
−
 ( + 1)
2
 ( + 1)(2 + 1)
6
−
 ( + 1)
2</p>
      <p>We can finally write after simplifications to define 
 = 
 
−
 ( + 1)
2

 (
12
− 1)</p>
      <sec id="sec-6-1">
        <title>If the value  is close to zero, the distance between IOs in time has no trend towards change.</title>
        <p>This suggests that if these IOs were the candidates to be identified as one object by separate values
 , then most likely these IOs really belong to the same physical object. If the value  is negative,
then the distance between the IOs decreases over time, and therefore, the possibility that the
identification problem is solved correctly increases. If the value  is significantly greater than zero,
most likely the IOs under consideration should not be identified as a single object, and the low
value</p>
        <p>at a particular point in time from the interval was random. Moreover, a larger value 
gives greater confidence that the IO data refers to different physical objects.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>8. Conclusions</title>
      <p>The paper proposes a method for solving the problem of identifying IOs that enter the monitoring
system’s common information resource from several data sources. For this purpose, it is proposed
to use a new proximity measure (similarity) of IO, which takes into account the nature of
uncertainty of the type "probability" for quantitative features and the type "possibility" for
qualitative features. At the same time, it does not require the transformation of feature values,
which significantly simplifies the formation of a metric – a proximity (or distance) function
between IOs as a whole according to all available features. The proposed measure was checked for
compliance with the mandatory conditions for the validity of measures.</p>
      <sec id="sec-7-1">
        <title>Performing the IO identification procedure allows the consumer to avoid duplication or data conflict, as well as increases the accuracy of IO representation in the monitoring system. Additionally, it is proposed to analyze the linear trend of the change in time of the distance 91</title>
        <p>between the IOs to be identified, which can be calculated using the least squares method, which
improves the quality of solving the identification problem.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <sec id="sec-8-1">
        <title>The author(s) have not employed any Generative AI tools.</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>I. Mandel</surname>
          </string-name>
          ,
          <article-title>Cluster analysis</article-title>
          ,
          <source>Finance and statistics</source>
          , Moscow,
          <year>1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Newcombe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kennedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Axford</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. James.</surname>
          </string-name>
          “
          <article-title>Automatic linkage of vital records</article-title>
          .
          <source>” Science</source>
          <volume>130</volume>
          .3381 (
          <year>1959</year>
          ):
          <fpage>954</fpage>
          -
          <lpage>959</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I. P.</given-names>
            <surname>Fellegi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Sunter</surname>
          </string-name>
          .
          <article-title>“A Theory for Record Linkage</article-title>
          .
          <source>” Journal of the American Statistical Association</source>
          <volume>64</volume>
          .328 (
          <year>1969</year>
          ):
          <fpage>1183</fpage>
          -
          <lpage>1210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <article-title>A unified probabilistic approach modeling relationships between attributes and objects</article-title>
          ,
          <source>in: Proceedings of the IEEE International Conference on Computer Vision</source>
          , Sydney, Australia,
          <year>2013</year>
          , pp.
          <fpage>2120</fpage>
          -
          <lpage>2127</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCV.
          <year>2013</year>
          .
          <article-title>4</article-title>
          . URL: https://openaccess.thecvf.com/content_iccv_2013/papers/Wang_A_Unified_
          <article-title>Probabilistic_2013 _ICCV_paper</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Nagarajan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Grauman</surname>
          </string-name>
          ,
          <article-title>Attributes as Operators: Factorizing Unseen Attribute-Object Compositions</article-title>
          ,
          <source>in: Proceedings of the European Conference on Computer Vision</source>
          ,
          <year>2018</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.
          <year>1803</year>
          .
          <volume>09851</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Singla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Domingos</surname>
          </string-name>
          ,
          <article-title>Object Identification with Attribute-Mediated Dependences</article-title>
          , in: A.M.∘
          <string-name>
            <surname>Jorge</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Torgo</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Brazdil</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Camacho</surname>
          </string-name>
          , J. Gama (Eds),
          <source>Knowledge Discovery in Databases: PKDD 2005, Lecture Notes in Computer Science()</source>
          , vol
          <volume>3721</volume>
          , Springer Berlin, Heidelberg,
          <year>2005</year>
          , pp.
          <fpage>297</fpage>
          -
          <lpage>308</lpage>
          . doi:
          <volume>10</volume>
          .1145/775047.775116. URL: https://alchemy.cs.washington.edu/papers/pdfs/singla-domingos05b.
          <fpage>pdf</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Richman</surname>
          </string-name>
          ,
          <article-title>Learning to match and cluster large high-dimensional data sets for data integration, in: Proceedings of the 8th ACM SIGKDD international conference on Knowledge discovery and data mining</article-title>
          ,
          <source>KDD '02</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York United States,
          <year>2002</year>
          , pp
          <fpage>475</fpage>
          -
          <lpage>480</lpage>
          . URL: https://www.cs.cmu.edu/~wcohen/postscript/kdd-2002.pdf
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarawagi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhamidipaty</surname>
          </string-name>
          ,
          <article-title>Interactive deduplication using active learning</article-title>
          ,
          <source>in: Proceedings of the 8th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          ,
          <source>KDD '02</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York,
          <year>2002</year>
          , pp
          <fpage>269</fpage>
          -
          <lpage>278</lpage>
          . URL: https://www.scribd.com/document/23688804/05-InteractiveDeduplicationUsingActiveLearning
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Pasula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Marthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Milch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Shpitser</surname>
          </string-name>
          ,
          <article-title>Identity uncertainty and citation matching</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>15</volume>
          (
          <year>2003</year>
          ):
          <fpage>1401</fpage>
          -
          <lpage>1408</lpage>
          . URL: https://people.csail.mit.edu/milch/papers/nipsnewer.pdf
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zadeh</surname>
          </string-name>
          ,
          <article-title>Fuzzy sets as a basis for a theory of possibility</article-title>
          ,
          <source>Fuzzy Sets and Systems 1.1</source>
          (
          <year>1978</year>
          ):
          <fpage>3</fpage>
          -
          <lpage>28</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0165</fpage>
          -
          <lpage>0114</lpage>
          (
          <issue>78</issue>
          )
          <fpage>90029</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dubois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Prade</surname>
          </string-name>
          ,
          <source>Possibility Theory: An</source>
          Approach to Computerized Processing of Uncertainty, Springer Science &amp; Business Media, USA,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Yu</surname>
            .
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Voronin</surname>
          </string-name>
          ,
          <article-title>Classification theory and its applications</article-title>
          , Nauka, Novosibirsk,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Mirkin</surname>
          </string-name>
          ,
          <article-title>Qualitative features and structures analysis</article-title>
          ,
          <source>Finance and Statistics</source>
          , Мoscow,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Yu. I. Zhuravlev</surname>
          </string-name>
          ,
          <article-title>On the application of algebraical techniques in pattern recognition and classification problems. Pattern recognition, classification, forecasting</article-title>
          , 1st. ed.,
          <source>Nauka</source>
          , Moscow,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Norvig</surname>
          </string-name>
          , Artificial Intelligence:
          <string-name>
            <given-names>A Modern</given-names>
            <surname>Approach</surname>
          </string-name>
          , 3rd ed.,
          <source>Pearson</source>
          ,
          <year>2009</year>
          . URL: https://api.pageplace.de/preview/DT0400.9781292153971_
          <issue>A27091185</issue>
          /preview9781292153971_A27091185.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>