<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Structure Preserving Exemplar-Based 3D Texture Synthesis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrew Babichev</string-name>
          <email>andrey.babichev@graphics.cs.msu.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Frolov</string-name>
          <email>vfrolov@graphics.cs.msu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Keldysh Institute of Applied Mathematics</institution>
          ,
          <addr-line>Miusskaya sq., 4, Moscow, 125047</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>GSP-1, Leninskie Gory, Moscow, 119991</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we propose exemplar-based 3D texture synthesis method which unlike existing neural network approaches preserve structural elements in texture. The proposed approach does this by accounting additional image properties which stand for the preservation of the structure with the help of a specially constructed error function used for training neural networks. Thanks to the proposed solution we can apply 2D texture to any 3D model (even without texture coordinates) by synthesizing high quality 3D texture and using local or world space position of surface instead 2D texture coordinates (fig. 1). Our solution is based on introducing 3 different error components in to the process of neural network fitting which helps to preserve desired properties of generated texture. The first component is for structuredness of the generated texture and the sample, the second component increases the diversity of the generated textures and the third one prevents abrupt transitions between individual pixels.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>3D texture synthesis, neural network, exemplar-based texture synthesis, structure preserving.</title>
      <sec id="sec-1-1">
        <title>1. Introduction</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Textures are one of the main components of realistic image synthesis. Exemplar based texture synthesis is used for generating new textures of a desired resolution which has similar appearance with the input texture but different pixels (like different parts of the road of the wall).</title>
      <p>model</p>
      <p>2021 Copyright for this paper by its authors.</p>
    </sec>
    <sec id="sec-3">
      <title>In this work we propose new method for exemplar-based 3D texture synthesis. A 3D texture can be thought of as a voxel grid, where each voxel contains a certain color and each plane cut in any direction of which is a regular 2D texture (fig. 1).</title>
      <sec id="sec-3-1">
        <title>2. Related work</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>When studying existing methods of texture synthesis it is needed to consider two principal problems:</title>
      <p>
        (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) how to define similarity between exemplar and generated images and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) how to generate the new
texture. Most existing methods for exemplar-based texture synthesis can be divided into two categories.
      </p>
    </sec>
    <sec id="sec-5">
      <title>The first group of methods defines a function of the properties of the exemplar texture, and then</title>
      <p>tries to modify a random noise texture until property functions of generated and exemplar images
become similar to the function of the properties of the exemplar texture. Either neural network based
approaches and pyramidal synthesis can be attributed to that group [15]. As the pronounced properties
of this group of methods, one can note high variability of the synthesized textures, low speed, as well
as in many cases poor results when working with textures that have high resolution or complex structure.</p>
    </sec>
    <sec id="sec-6">
      <title>This methods ignore large details and poorly preserve structural features. Some of these shortcomings</title>
      <p>have been eliminated in recent works by consistently doubling the resolution of the generated textures
generate the new texture. Typical representatives of this group of methods are per-pixel [16] and
perpatch synthesis [17, 18]. As a characteristic of this group, one can designate the high speed and
acceptable results on texture sampling. The disadvantage of this group of methods is the low variability
of the synthesized textures.
2.1.</p>
      <sec id="sec-6-1">
        <title>Neural network sythesis</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Most existing neural-network based texture synthesis methods [1-5] uses VGG-19 [9]. Neural</title>
      <p>network generators first define the image property function. For this, an exemplar of texture is fed into
the neural network, and the activation function is calculated for each layer  of the neural network. Each
activation function generates a certain set of filtered images - the so-called feature maps. Each layer
with   filers will have same amount of feature maps, each of which has a total dimension   .
Therefore, all feature maps can be stored in matrix   ∈    ×  , where    – is an activation for j filter
at spatial position  inside layer  .</p>
      <p>After finding all feature maps, for each of them we can calculate the Gram matrix   ∈    ×  :
A set of Gram matrices { 1,  2. . . ,   } of neural network layers 1, 2, …, L is a possible way to
describe the properties of an image. Then, to define the similarity function of the exemplar texture and
the synthesized texture, we can use the following error   :
   = ∑       .</p>
      <p>1
  ( ,  ) = ∑      

‖  ( ) −   ( )‖ ,

2
neural network to the error   , ‖
‖</p>
      <p>– Frobenius norm.
where  и  – generated texture and source texture respectively;   – contribution of each layer  of</p>
      <p>
        Thus, to synthesize a new texture, we need to minimize the specified error   between exemplar
texture and the random noise texture. This can be achieved by applying gradient descent to the noise
texture while calculating the gradients using backpropagation of errors.
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <sec id="sec-7-1">
        <title>2.2. 3D texture synthesis</title>
      </sec>
      <sec id="sec-7-2">
        <title>2.2.1. 3D texture generator concept</title>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>A similar approach to the two-dimensional case is used in the synthesis of three-dimensional textures</title>
      <p>[10-14]. The algorithm can be described as follows: we have a certain texture generator based on
convolutional neural networks and which depends on a fixed set of parameters  . Three-dimensional
white noise of several different resolutions  is fed to this generator, and the resolution of the
synthesized texture will depend on the resolution of the supplied noise. Passing noise through itself, the
generator will produce a three-dimensional texture of the specified resolution.</p>
    </sec>
    <sec id="sec-9">
      <title>During training the generated texture will be divided into all possible planes along the three</title>
      <p>directions of the coordinate axes, and using some function that describes the properties of the image, it
will calculate the error between the planes representing generated 3D texture and the exemplar textures
we have. Setting the necessary parameters of the generator  to generate a high-quality texture will be
done using gradient descent and back propagation of the error similar to the two-dimensional case.</p>
      <sec id="sec-9-1">
        <title>2.2.2. Generator architecture</title>
        <p>The generator architecture is shown in fig. 2. It is a sequence of convolutional blocks, upscaling
blocks, and concatenation blocks:
 A convolution block is a sequence of 3 convolutional layers, the first two of which have 3x3x3
kernels, and the last - 1x1x1. This is followed by the batch normalization layer [20] and the Relu
activation function [21]. Thus, after this layer, the size of the texture is reduced by 4.
 The upscale block doubles 3D texture resolution (each voxel will be copied 8 times).
 The concatenation block performs batch normalization and then concatenates our textures into
channels. If the textures are Are different in size then they are cropped to a smallest size.</p>
      </sec>
      <sec id="sec-9-2">
        <title>2.2.3. Parameters fitting</title>
        <p>As already mentioned, to calculate the similarity between the three-dimensional texture obtained by
the generator and the exemplar texture we have, the first one will be divided into all possible planes in
all directions of the coordinate axes. Then the error can be written as:</p>
        <p>−1
1
 =1    =0
 ( ,  ) =
∑</p>
        <p>∑  2(  , ,  ) ;
 2(  , ,  ) =   (  , ,  ),
where  – means axis (x, y or z) in which 2D texture was split,   – number of planes for axis d in
which the texture was split,   , –  -th plane in the d direction of the coordinate axis of the synthesized
three-dimensional texture and  2 – is an error implying the difference between two-dimensional texture
exemplars. For calculation of activation maps that are required by our error function, we use the same
neural network as in VGG-19 paper.</p>
      </sec>
      <sec id="sec-9-3">
        <title>3. Proposed method</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Using just a single Gram matrix to calculate the similarity between two-dimensional samples of</title>
      <p>textures is not an optimal approach since the Gram matrix itself poorly represents such aspects of the
texture as the presence of structure, smooth inter-pixel transitions, and many others. For this reason, the
existing exemplar-based methods [10, 11] of three-dimensional synthesis poorly work for input textures
with complex structure: the resulting texture looks “broken”.</p>
      <p>To solve these problems of the two-dimensional synthesis, in [2] it was proposed to generate a
texture only not using the Gram matrix, but also a combination of several error components; moreover,
an approach that allows preserving the structure of synthesized textures was presented in [3, 4, 8]. We
have further developed idea from [2] to solve 3D synthesis problems and added three additional error
components to train the generator proposed in [10]: using the first component, we have compared the
structuredness of the generated texture and the sample. In fact, this error is the autocorrelation of feature
maps of the given layers of the neural network, multiplied by the coefficient:
   , =</p>
      <p>1
(  − | |)(  − | |)
;  ∈ [
2</p>
      <p>2
−  ,   ] ;  ∈ [
−   
where   и   – spatial sizes of feature maps   (  ∗   =  
. Also for coefficient   Gauss window
could be used. The second introduced component of the error function increases the diversity of the
generated textures and is the usual difference between feature maps:
 ,</p>
      <p>1
   
  (  , ,  ) = ∑</p>
      <p>‖  (  , ) −   ( )‖ ;

  (  , ,  ) = ∑‖  (  , ) −   ( )‖ .</p>
      <p>(  , ,  ) = ∑‖  (  , ) −   ( )‖ ;

(9)
(10)
(11)</p>
    </sec>
    <sec id="sec-11">
      <title>The third component prevents abrupt transitions between individual pixels and is some kind of antialiasing: given in Table 1.</title>
      <p>where  – is a parameter of and algorithm.</p>
    </sec>
    <sec id="sec-12">
      <title>Therefore, final difference between two 2D exemplars is:</title>
      <p>Parameter values, used by us for texture synthesis</p>
    </sec>
    <sec id="sec-13">
      <title>Layer names</title>
    </sec>
    <sec id="sec-14">
      <title>Relu11, Relu21,</title>
    </sec>
    <sec id="sec-15">
      <title>Relu31, Relu41,</title>
    </sec>
    <sec id="sec-16">
      <title>Relu51</title>
    </sec>
    <sec id="sec-17">
      <title>Pool2</title>
    </sec>
    <sec id="sec-18">
      <title>Pool2</title>
    </sec>
    <sec id="sec-19">
      <title>Relu11</title>
    </sec>
    <sec id="sec-20">
      <title>Layer weights</title>
      <p>0.2, 0.2, 0.2, 0.2,
0.2
1
1
1</p>
    </sec>
    <sec id="sec-21">
      <title>Loss weight</title>
      <p>=0.5
 =0.5*1e-6
 =-1*1e-4
 =-1*1e-3</p>
    </sec>
    <sec id="sec-22">
      <title>Miscellaneous</title>
      <p>=1e-3</p>
      <sec id="sec-22-1">
        <title>4. Experimental evaluation and comparison</title>
      </sec>
    </sec>
    <sec id="sec-23">
      <title>To test and compare the proposed method, the following two groups of textures were used:</title>
      <p>1.</p>
      <p>Textures were selected in the first group (Fig. 3,4) in order to find out how high-quality and
original the textures synthesized by the generator, and also whether a proposed error components
interferes with the synthesis. The textures were chosen, consisting of a monochrome background
with some chaotic picture on top of it. These textures were chosen with the idea that the generator
should work reasonably on this type of textures due to its simplicity, but the result may be too similar
to the original texture or differ in random outbursts of colors atypical for this texture.</p>
    </sec>
    <sec id="sec-24">
      <title>Textures were selected in the second group (Fig. 5,6) in order to test the generator's ability to preserve structure. We take textures with repeating patterns with a small number of colors used and without sharp gradient transitions. Such textures were chosen in order to exclude the influence on the synthesis of characteristics not related to the structural organization of the texture.</title>
    </sec>
    <sec id="sec-25">
      <title>Finally, we have performed visual comparison of our method to [11] at figures 7 and 8.</title>
    </sec>
    <sec id="sec-26">
      <title>Unfortunately, we were not able to run their implementation due to version conflicts of libraries, either we didn’t find original texture in desired resolution. Therefore, our comparison is not strict but shows that both methods preserve details well.</title>
      <sec id="sec-26-1">
        <title>5. Conclusions</title>
        <p>It can be seen from fig. 3-6 that the proposed method was able to show better results than it is base
version [10]. Like [10], it works well with textures from the first group - the images obtained by the
generator are new textures without any color outliers. Additional error components not only did
preserve the stability of the synthesis, but also helped to broaden the already wide variety of generated
textures, as well as improved the sharpness of pixel transitions (fig. 3, 4). If we look at the results of the
second group, it can be seen that the proposed method preserves the structure better in textures
containing structural patterns (fig. 5, 6). Thus, introducing additional components to the error helped to
overcome the fundamental problem of poor preservation of structuredness inherent to the group of
methods to which neural network synthesis belongs, and in the problem of synthesizing
threedimensional textures.</p>
      </sec>
      <sec id="sec-26-2">
        <title>6. References</title>
        <p>[7] R. P éteri, S. Fazekas, M. J. Huiskes, DynTex: A Comprehensive Database of Dynamic Textures,
2010. URL: http://dyntex.univ-lr.fr/database.html
[8] Y. Zhou, Z. Zhu, X. Bai, D. Lischinski, D. Cohen-Or, H. Huang, Non-Stationary Texture Synthesis
by Adversarial Expansion, arXiv preprint (2018) arXiv:1805.04487
[9] A. Frühstück, I. Alhashim, P. Wonka, TileGAN: synthesis of large-scale non-homogeneous textures.</p>
        <p>
          ACM Trans. Graph. 38(
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) (2019) Article 58. doi: 10.1145/3306346.3322993
[10] J. Gutierrez, J. Rabin, B. Galerne, T. Hurtut, On Demand Solid Texture Synthesis Using Deep 3D
        </p>
      </sec>
    </sec>
    <sec id="sec-27">
      <title>Networks, Computer Graphics Forum 36 (2019). doi:10.1111/cgf.13889</title>
      <p>[11] P. Henzler, N. J. Mitra, T. Ritschel, Learning a Neural 3D Texture Space From 2D Exemplars, in:</p>
    </sec>
    <sec id="sec-28">
      <title>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),</title>
      <p>2020, pp. 8356–8364.
[12] D. J. Rezende, S. Eslami, S. Mohamed, P. Battaglia, M. Jaderberg, N. Heess, Unsupervised
Learning of 3D Structure from Images, Advances in neural information processing systems 29
(2016) 4996–5004.
[13] J. Gwak, C. B. Choy, M. Chandraker, A. Garg, S. Savarese, Weakly supervised 3d reconstruction
with adversarial constraint, in: International Conference on 3D Vision, 2017, pp. 263–272.
doi:10.1109/3DV.2017.00038
[14] X. Yan, J. Yang, E. Yumer, Y. Guo, H. Lee, Perspective Transformer Nets: Learning Single-View</p>
    </sec>
    <sec id="sec-29">
      <title>3D Object Reconstruction without 3D Supervision, arXiv preprint (2016) arXiv:1612.00814.</title>
      <p>[15] J. S. De Bonet, Multiresolution sampling procedure for analysis and synthesis of texture images,
in: Proceedings of the 24th annual conference on Computer graphics and interactive techniques.</p>
    </sec>
    <sec id="sec-30">
      <title>ACM Press/Addison-Wesley Publishing Co., New York, NY, USA, 1997, pp. 361-368.</title>
      <p>
        doi:10.1145/258734.258882
[16] A. A. Efros, T. K. Leung, Texture Synthesis by Non-Parametric Sampling, in: Proceedings of the
seventh IEEE international conference on computer vision, 1999, pp. 1033–1038.
[17] A. A. Efros, W. T. Freeman, Image quilting for texture synthesis and transfer, in: Proceedings of
the 28th annual conference on Computer graphics and interactive techniques. ACM, New York, NY,
USA, 2001, pp. 341-346. doi: 10.1145/383259.383296
[18] L. Liang, C. Liu, Y.-Q. Xu, B. Guo, H.-Y. Shum, Real-time texture synthesis by patch-based
sampling, ACM Trans. Graph. 20(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) (2001) 127–150. doi:10.1145/501786.501787
[19] K. Simonyan, A. Zisserman, Very Deep Convolutional Networks for Large-Scale Image
      </p>
    </sec>
    <sec id="sec-31">
      <title>Recognition, arXiv preprint (2014) arXiv:1409.1556.</title>
      <p>[20] S. Ioffe, C. Szegedy, Batch normalization: accelerating deep network training by reducing internal
covariate shift, in: Proceedings of the 32nd International Conference on International Conference on
Machine Learning, 2015, pp. 448–456.
[21] X. Glorot, A. Bordes, Y. Bengio, Deep Sparse Rectifier Neural Networks, in: Proceedings of the
fourteenth international conference on artificial intelligence and statistics, JMLR Workshop and</p>
    </sec>
    <sec id="sec-32">
      <title>Conference Proceedings, 2011, pp. 315–323.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Gatys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Ecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bethge</surname>
          </string-name>
          ,
          <article-title>Texture synthesis using convolutional neural networks</article-title>
          , in: C.
          <string-name>
            <surname>Cortes</surname>
            ,
            <given-names>D. D.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sugiyama</surname>
          </string-name>
          , and R. Garnett (Eds.),
          <source>Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1 (NIPS'15)</source>
          , Vol.
          <volume>1</volume>
          . MIT Press, Cambridge, MA, USA,
          <year>2015</year>
          , pp.
          <fpage>262</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O.</given-names>
            <surname>Sendik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cohen-Or</surname>
          </string-name>
          ,
          <article-title>Deep Correlations for Texture Synthesis</article-title>
          .
          <source>ACM Trans. Graphics</source>
          <volume>36</volume>
          (
          <issue>5</issue>
          ) (
          <year>2017</year>
          )
          <article-title>Article 161</article-title>
          . doi:
          <volume>10</volume>
          .1145/3015461
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wand</surname>
          </string-name>
          ,
          <article-title>Combining Markov Random Fields and Convolutional Neural Networks for Image Synthesis</article-title>
          ,
          <source>in: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>2479</fpage>
          -
          <lpage>2486</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gousseau</surname>
          </string-name>
          , G.-S. Xia,
          <article-title>Texture synthesis through convolutional neural networks and spectrum constraints</article-title>
          ,
          <source>in: Proceings of the 23rd International Conference on Pattern Recognition (ICPR)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>3234</fpage>
          -
          <lpage>3239</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICPR.
          <year>2016</year>
          .
          <volume>7900133</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ulyanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vedaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lempitsky</surname>
          </string-name>
          , Improved Texture Networks:
          <article-title>Maximizing Quality and Diversity in Feed-forward Stylization and Texture Synthesis</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>6924</fpage>
          -
          <lpage>6932</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tesfaldet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Brubaker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Derpanis</surname>
          </string-name>
          ,
          <article-title>Two-Stream Convolutional Networks for Dynamic Texture Synthesis</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>6703</fpage>
          -
          <lpage>6712</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>