<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Segmenting out Generic Objects in Monocular Videos</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Hula</string-name>
          <email>jan.hula@osu.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Adamczyk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Mojzisek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vojtech Molek</string-name>
          <email>vojtech.molek@osu.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CE IT4I - IRAFM, University of Ostrava 30. dubna 22</institution>
          ,
          <addr-line>701 03 Ostrava</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an approach for generic object detection and segmentation in monocular videos. In this task, we want to segment objects from a background with no prior knowledge about the possible classes of objects which we may encounter. This makes this task much harder than the classical object detection and segmentation, which can be posed as a supervised learning problem. Our approach uses an ensemble of 3 different models which are trained by different objectives and have different failure modes and therefore complement each other. We demonstrate the usefulness of our approach on a custom dataset containing 18 classes of organic objects. Using our method, we were able to recover the classes of objects in a fully unsupervised way.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Separating generic objects from a background in
monocular videos is a challenging task. We believe that this
problem is essential to Computer Vision and, as such, gained
an unproportionally small amount of attention from the
research community. The ability to separate objects from the
background would vastly simplify other tasks as it can be
viewed as a kind of dimensionality reduction on relevant
features.</p>
      <p>In image classification, object separation prevents a
classifier from learning spurious correlations, which could
arise when a certain class is often captured on a particular
background. Object separation from a background
automatically restricts a classifier to consider only features
directly tied to the class label, as opposed to features only
correlated with it.</p>
      <p>As the generic object separation is a rather nonstandard
task and still vaguely defined – it is not clear what should
be considered an independent object – we focus on a
simplified setting in which a camera captures a single salient
object. Our solution is an ensemble of three different
models trained for different objectives.</p>
      <p>Furthermore, we study the impact of the background
removal on the clustering properties of the resulting
representations. The representations are obtained by training
a neural network in an unsupervised way. We are able
to recover the categories of objects in a fully unsupervised
way, using our custom dataset containing videos of organic
things.</p>
      <p>Our main contributions are:
• We present an ensemble model that can separate
objects from the background in monocular videos
containing one salient object. The ensembled models
compensate for each other failure modes.
• We demonstrate the benefits of object separation by
comparing the classification accuracy of objects with
and without background using clustering.</p>
      <p>Section 2 contains related work. Section 3 describes
our approach for detecting and segmenting generic objects
within monocular videos. Section 4 describes how the
detected objects enable unsupervised discovery of object
classes. In section 5 we describe our experiments and the
dataset we test our approach on, and finally we provide a
conclusion in section 6.
move independently. Nonetheless, we consider this as our
working definition because it allows us to make progress
in generic object detection and segmentation.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Generic object separation is a largely an unexplored area,
and therefore similar works are scarce. Our approach
is most related to DINO method [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] which uses
selfsupervised learning with Vision Transformers. The
authors introduced DINO as a form of self-distillation with
no labels. They emphasize that DINO automatically learns
an interpretable representation and separates the main
object from the background clutter.
      </p>
      <p>
        Lu et al. introduced an approach called CO-attention
Siamese Network (COSNet) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for unsupervised video
object segmentation. It is based on two ideas. The first
is the importance of inherent correlation among video
frames, and the second is the global co-attention
mechanism responsible for learning motion in short-term
temporal segments. The COSNet is trained on pairs of video
frames, which increases the learning capacity.
      </p>
      <p>
        The task of class discovery is marginally related to
current self-supervised approaches using large amounts of
unlabeled data such as [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] and approaches that try to
exploit coherency in the data [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Lastly, our approach for class discovery can be seen as a
version of clustering with constraints. It has been heavily
studied in the past, for example, by [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Generic Object Detection and</title>
    </sec>
    <sec id="sec-4">
      <title>Segmentation</title>
      <p>This section describes our approach to generic object
detection and segmentation. By a generic object we mean
an object of an unknown class. We use this term to
distinguish it from classical object detection and
segmentation, which can deal only with a concrete set of
specified classes. Classical object detection and segmentation
is much easier because it can be approached as a
supervised learning problem on a dataset with annotated
bounding boxes and segmentation masks. With generic objects,
it is not that straightforward because it is not known in
advance what kind of objects we will encounter at the test
time.</p>
      <p>Moreover, at first it may not be obvious how to define
what should be considered as a separate object. One
useful definition would be that an object is anything that can
move independently from the rest of the environment. In
this view, we can understand generic object segmentation
as a way to factorise the visual stream into independent
components. We need to mention that this definition does
not cover all cases in which we would like to detect
something as a separate object. Examples include buildings,
letters on a sheet of paper, and other “entities” which can not
3.1</p>
      <sec id="sec-4-1">
        <title>Ensemble of Models Trained for Different</title>
      </sec>
      <sec id="sec-4-2">
        <title>Objectives</title>
        <p>Our approach to the problem of generic object detection
and segmentation is based on an ensemble of 3 models
which are trained by different objectives. Even though
each of these models has its own failure modes, together
they constitute a robust ensemble. Concretely, we use
one model trained for depth map prediction, one model
trained for optical flow estimation, and one model trained
for tracking of objects. Using the model for depth
prediction, we can separate foreground objects based on depth,
using the model trained for optical flow estimation, we
can separate moving objects, and finally, using the tracker,
we can verify the temporal coherency of our predictions.
The tracker is initialised with a bounding box obtained
from the predictions of the two other models in the frames
where these predictions are most consistent. The
following paragraphs provide a high-level description of these
three models. For a more complete description of these
models, see the respective publications.</p>
        <p>
          Depth Prediction For the depth prediction, we use the
model introduced by Ranftl et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], available from the
author’s repository1. This transformer-based model
predicts a scalar value for each pixel, which represents the
distance of the surface captured by that pixel from the camera
center.
        </p>
        <p>
          Optical Flow Estimation For optical estimation, we use
the model introduced by Teed et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. It is also
a transformer-based model which requires 2 consecutive
frames of video to produce the optical flow field. The
optical flow field assigns two scalar values to each pixel.
These values represent the pixel displacement on the x and
y axes, relative to the previous frame. To obtain one scalar
value for each pixel, we take the magnitude of the
displacement. We used the implementation of the model with
trained weights provided in the authors repository2.
Object tracking For tracking objects, we use a model
called SiamMask [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a neural network trained as a
Siamese architecture which simultaneously performs both
visual object tracking and object segmentation in a video.
We used the implementation available online3.
1https://github.com/intel-isl/DPT
2https://github.com/princeton-vl/RAFT
3https://github.com/foolwood/SiamMask
In our dataset, each video captures a hand holding one
object. Using our approach for generic object segmentation
which is based on predicted depth maps and optical flow,
our model segments out the hand together with the object.
We fix this issue by segmenting out hands separately by a
model trained specifically for hand segmentation.
        </p>
        <p>
          To obtain training data for hand segmentation, we
downloaded the following 4 datasets: GTEA,
HandOverFace, GTEA_GAZE_PLUS, and EgoHands456. The
architecture of the model is UNet with timm_regnetty_160
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] as encoder and softmax2D as activation. The encoder
weights were pretrained on ImageNet.
        </p>
        <p>Using freely available datasets for hand segmentation
mentioned above, the trained model was not working well
on our dataset, probably because of a large distribution
shift (most of the images in these public datasets contained
hands in front of the face or were captured in the interior).</p>
        <p>To mitigate this problem, we used a simple trick to
enlarge the training data with images of hands which are
similar to the images in our target dataset. Concretely, we
captured our hands from a similar viewpoint as in our dataset,
and then used the same model for depth prediction to
produce depth maps for every 10th frame within the video.
Finally, we thresholded the predicted death maps to obtain
reliable segmentation masks of hands. In this way, we
obtained hundreds of labeled images of hands with minimal
effort. After adding this dataset to the other datasets, we
obtained an accurate model for hand segmentation.</p>
        <p>We use this model to remove hands from the mask
predicted by our ensemble. More precisely, we remove the
hands from the outputs of the model predicting the depth
map and the model predicting the optical flow before we
initialise the bounding box for the tracker. In this way we
obtain masks only for the object, ignoring the hands.
3.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Bounding Box Initialization</title>
        <p>As mentioned above, we initialize the tracker with a
bounding box obtained from the predictions of depth and
optical flow maps within each frame. These predictions
are two rectangular matrices of the same shape. We rescale
the range of values to the interval between 0 and 1 and
denote the final matrices by f1(x) and f2(x), respectively.</p>
        <p>
          We first choose k frames where the predictions from
these two models are most consistent. To measure this
consistency, we devise the following heuristic. We first
detect edges using a Canny edge detector [
          <xref ref-type="bibr" rid="ref13">12</xref>
          ] in both
predictions and then measure the overlap of the resulting edges.
To account for small deviations of edges in the two
predictions, we blur them with a gaussian kernel of the width
set to 7px to achieve their overlap if they are close to each
4http://cbs.ic.gatech.edu/fpv/
5https://www.cl.cam.ac.uk/research/rainbow/emotions/hand.html
6http://vision.soic.indiana.edu/projects/egohands/
other. Then we compute the consistency score c of frame
x using the following formula:
c(x) =
        </p>
        <p>å
i2Pixels(x)
ee1i+e2i ;
(1)
where e1 and e2 are the two blurred edge maps from the
two predictions and i indexes individual pixels. Next, we
obtain the aggregated predictions for each pixel i by:
y(x)i = e f1(x)i+ f2(x)i +
1</p>
        <p>å
jPixels(x)j j2Pixels(x)
e f1(x) j+ f2(x) j :
(2)
We exponentiate the sum of the predictions from the two
models because we want these predictions to interact
superlinearly. We also subtract the mean of this value taken
over all pixels within the image to make the aggregated
predictions centered at zero. Therefore, the pixels where
no object was predicted will contain negative values.</p>
        <p>Once we have the aggregated predictions for the
selected frames, we initialize the bounding boxes for the
tracker. For this, we again devise a score which captures
how well a given bounding box (bbox) covers pixels with
high values (signifying that an object is present) and at the
same time excludes pixels with low values. It has the
following form:</p>
        <p>å
i2Pixels(x)
bboxScore(bbox; y(x)) =
isInBbox(i; bbox) y(x)i;
(3)
where the function isInBbox returns 1 if the pixel is
not contained in the bounding box and 1 otherwise.</p>
        <p>
          Finally, we optimize the coordinates of the bounding
box using CMA-ES [
          <xref ref-type="bibr" rid="ref14">13</xref>
          ] which is a derivative-free
optimization algorithm used for the optimization of
continuous parameters. The optimization tries to find coordinates
which maximize this score. At the end of this procedure,
we obtain k frames with bounding boxes in each video.
The quality of predictions and the resulting bounding box
is shown in Figure 2.
Using the chosen frames and their bounding boxes, we
initialize the tracker and let it track the object in between
the selected frames. The tracker provides another layer of
consistency check. Once we have the predictions from the
three models (denoted by f1(x), f2(x) and f3(x)), we can
treat the consistency of these predictions as a certainty of
the whole ensemble. To measure this certainty, we again
compute the consistency score as we did in the selection
of reliable frames in the Equation 1. Using an empirically
estimated threshold, we filter out frames with low
consistency and for each pixel i in the filtered frames, we
aggregate the predictions of the ensemble with the following
formula:
output(x)i =
min eå3j=1 f j(x)i
1; e2
        </p>
        <p>1
e2
1
;
(4)
The subtraction of 1 ensures that we get 0 when all three
models predict 0. Thresholding and dividing by e2 1
insures that we obtain a value close to 1 when at least
two models predict values close to 1. Finally, we obtain a
bounding box for each frame using the same method as in
the bounding box initialization (optimization using
CMAES), i.e., minimizing the objective in Equation 1.</p>
        <p>The whole process can be viewed as certainty
propagation. We first select a few frames where the first two
models agree on their predictions and from these the tracker
propagates the certainty to other frames.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Evaluating The Quality of the Aggregated Predictions</title>
        <p>Our final task is a discovery of classes of objects within
monocular videos. We view the generic object
detection and segmentation as an intermediate step towards this
goal. Therefore, we evaluate the usefulness of our
approach on this target task. That is, we test how well are
we able to recover classes of objects from images where
the objects were segmented out by our approach. We
compare it to the setup where we use the same algorithm for
class discovery but where we use the original images with
a background. We also mention that we do not require
pixel-perfect segmentation masks as our goal is only to
focus on the relevant parts of the image, so that the measured
similarity between images will mostly reflect the
similarity of objects and not of backgrounds. The next section
describes our pipeline for the task of class discovery
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Class Discovery</title>
      <p>
        The algorithm for class discovery was proposed in [
        <xref ref-type="bibr" rid="ref15">14</xref>
        ].
Its input is a set of videos, each containing one object and
each represented as a sequence of images. The goal is
to find clusters of videos based on the similarity between
them. Generally, our algorithm works in 3 steps:
1. Measure the similarity between every pair of videos
with the method described in Section 4.1.
2. Construct a similarity graph by connecting each
video to its five most similar videos.
3. Apply the Louvain community detection
algorithm [
        <xref ref-type="bibr" rid="ref16">15</xref>
        ] to detect the highly interconnected parts
of the graph and consider these as the discovered
classes.
      </p>
      <p>The advantage of the Louvain algorithm is that it needs no
apriori knowledge of the number of clusters/communities.</p>
      <p>The accuracy of our approach is measured in two ways.
First, by counting how many times a video was assigned
to an incorrect cluster. Second, whether the algorithm
discovered all clusters. It is clear that the final accuracy
mostly reflects the measured similarity between individual
videos.
Each video is represented by a sequence of images, but to
compute the similarity, we ignore the ordering and treat the
sequence as a set. The similarity between the two videos
is computed in the following four steps:
1. Train an autoencoder using images from all videos
to obtain a low-dimensional representation zi of each
image xi.
2. In each video, select n representative frames which
are not correlated, described in 4.2.
3. For each pair of videos, compute all pairwise
similarities d(zi; z j) with cosine distance.
4. Finally, select the l most similar pairs of images and
average their similarities to obtain the final similarity
between two videos.</p>
      <p>The intuition behind step 4 is that videos of similar
objects may contain only a few frames where these objects
are captured from the same angle or in the same situation.
4.2</p>
      <sec id="sec-5-1">
        <title>Filtering out Correlated Frames</title>
        <p>Step 2 of similarity computation takes n representative
frames. If we would simply use all frames from each
video, the distribution of the dataset may end up skewed
because some parts of a video may be more static than
others. These static parts would produce many correlated
frames. Therefore, the correlated frames need to be filtered
out from a given video. We first test whether the
subjective visual similarity of images can be captured by cosine
similarity between their low-dimensional representations
obtained in step 1 of the similarity computation. As can
be seen in Figure 4, it captures the visual similarity well
enough.</p>
        <p>To extract n uncorrelated frames from each video,
we run k-means clustering, where k = n, on the
lowdimensional representations and take the most similar
frame to every centroid of the resulting clusters. This
simple heuristic produces uncorrelated images.</p>
        <p>To conclude this section, if the low-dimensional
representation of individual images obtained by the
autoencoder reflects the similarity between the captured objects
and not some other irrelevant factors, we may expect to
obtain meaningful clusters. Moreover, note the benefit of
creating the similarity graph of videos instead of
individual images. All images in one video are automatically
linked together. If a few images in 2 videos are similar, this
similarity is propagated to other frames within the video,
To test the algorithm for object discovery, we assembled
a custom dataset of organic objects. The dataset
contains 18 classes of organic objects, some of which are
depicted in Figure 5. We have chosen organic objects
because they naturally produce large variability between
instances. For every class, we capture ten different samples
from different viewpoints. The final dataset can be
downloaded at the following address – github.com/Jan21/
Organic-objects-dataset.</p>
        <p>Using our ensemble described in Section 3, we segment
the object in every frame of each video. Using the
resulting bounding boxes and segmentation masks, we crop
each image and blur the background to suppress the
distinctive features present in the background.</p>
        <p>To obtain the low-dimensional representations used to
filter out correlated frames and construct the similarity
graph, we resize all images to a fixed resolution of 64 64
pixels and train a convolutional autoencoder. The
autoencoder has five convolutional layers (16, 32, 64, 128, 256
filters with stride 2) and one fully-connected layer7 with
dimensions 1024 ! 96. 8
After running community detection on the similarity
graph, we inspected how many image bundles were
assigned to the wrong component. Out of 173 videos, only 5
in the training set were assigned to the wrong component.
The result of community detection on the constructed
similarity graph is shown in Figure 6. Numerical results and
comparison with clustering of non-segmented images are
presented in Table 1. From the accuracy and number of
discovered classes, it is clear that background removal
created a significant accuracy difference of 56:1% and a
difference in the number of correctly discovered classes.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this contribution, we present a method for generic object
detection and segmentation, which uses an ensemble of
7Image is downscaled to 2 2 times 256 filters ! 1024 input vector
to the fully-connected layer.</p>
      <p>8We also tried to extract representations by using VGG16 which was
pre-trained on ImageNet. These representations better discriminated very
similar objects (e.g., two types of red flowers).
three different models trained by three different objectives.
To demonstrate the effectiveness of the approach, we have
created a custom dataset of organic objects. The dataset
was used in our pipeline to remove background, create
low-dimensional representations, and perform class
discovery and classification. We have shown that background
removal significantly increases the accuracy of the
classification and the number of correctly discovered classes.</p>
      <p>In future work, we plan to optimize our approach for
speed, generalize it to work with videos containing
multiple objects, and make the class discovery an online process
that can discover new classes on-the-fly.</p>
      <p>Pre-segmenting the objects
Yes
No
Number of discovered classes
18
15</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Caron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          , I. Misra,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jégou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mairal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , “
          <article-title>Emerging properties in self-supervised vision transformers</article-title>
          ,
          <source>” arXiv preprint arXiv:2104.14294</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Porikli</surname>
          </string-name>
          , “
          <article-title>See more, know more: Unsupervised video object segmentation with co-attention siamese networks</article-title>
          ,
          <source>” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>3623</fpage>
          -
          <lpage>3632</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xie</surname>
          </string-name>
          , E. Hovy, M.-
          <string-name>
            <given-names>T.</given-names>
            <surname>Luong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          , “
          <article-title>Self-training with noisy student improves imagenet classification</article-title>
          ,” arXiv preprint arXiv:
          <year>1911</year>
          .04252,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Bagherinezhad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Horton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rastegari</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          , “
          <article-title>Label refinery: Improving imagenet classification through label progression,” ArXiv</article-title>
          , vol. abs/
          <year>1805</year>
          .02641,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhattacharjee and S. Das</surname>
          </string-name>
          , “
          <article-title>Temporal coherency based criteria for predicting video frames using deep multi-stage generative adversarial networks</article-title>
          ,
          <source>” in Advances in Neural Information Processing Systems</source>
          , pp.
          <fpage>4268</fpage>
          -
          <lpage>4277</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>I.</given-names>
            <surname>Davidson</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ravi</surname>
          </string-name>
          , “
          <article-title>Clustering with constraints: Feasibility issues and the k-means algorithm</article-title>
          ,”
          <source>in Proceedings of the 2005 SIAM international conference on data mining</source>
          , pp.
          <fpage>138</fpage>
          -
          <lpage>149</lpage>
          , SIAM,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bilenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Mooney</surname>
          </string-name>
          , “
          <article-title>Probabilistic semi-supervised clustering with constraints,” Semi-supervised learning</article-title>
          , pp.
          <fpage>71</fpage>
          -
          <lpage>98</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ranftl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bochkovskiy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Koltun</surname>
          </string-name>
          , “
          <article-title>Vision transformers for dense prediction</article-title>
          ,
          <source>” arXiv preprint arXiv:2103.13413</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Teed</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          , “Raft:
          <article-title>Recurrent all-pairs field transforms for optical flow</article-title>
          ,
          <source>” in European Conference on Computer Vision</source>
          , pp.
          <fpage>402</fpage>
          -
          <lpage>419</lpage>
          , Springer,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Bertinetto,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Torr</surname>
          </string-name>
          , “
          <article-title>Fast online object tracking and segmentation: A unifying approach</article-title>
          ,”
          <source>in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pp.
          <fpage>1328</fpage>
          -
          <lpage>1338</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Yakubovskiy</surname>
          </string-name>
          , “
          <article-title>Segmentation models pytorch</article-title>
          .” https://github.com/qubvel/segmentation_ models.pytorch,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Accuracy</surname>
          </string-name>
          <year>97</year>
          .1%
          <issue>41</issue>
          .0%
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Canny</surname>
          </string-name>
          , “
          <article-title>A computational approach to edge detection,” IEEE Transactions on pattern analysis and machine intelligence</article-title>
          ,
          <source>no. 6</source>
          , pp.
          <fpage>679</fpage>
          -
          <lpage>698</lpage>
          ,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          , “
          <article-title>The cma evolution strategy: a comparing review,” Towards a new evolutionary computation</article-title>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>102</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hula</surname>
          </string-name>
          , “
          <article-title>Unsupervised object-aware learning from videos</article-title>
          ,” in
          <source>2020 IEEE Third International Conference on Data Stream Mining &amp; Processing (DSMP)</source>
          , pp.
          <fpage>237</fpage>
          -
          <lpage>242</lpage>
          , IEEE,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>V. D.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Guillaume</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lambiotte</surname>
          </string-name>
          , and E. Lefebvre, “
          <article-title>Fast unfolding of communities in large networks</article-title>
          ,
          <source>” Journal of statistical mechanics: theory and experiment</source>
          , vol.
          <year>2008</year>
          , no.
          <issue>10</issue>
          , p.
          <fpage>P10008</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>