<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Interdisciplinary Classi cation of Audio E ects in the Audio E ect Ontology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Wilmering</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark B. Sandler</string-name>
          <email>mark.sandlerg@qmul.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Digital Music (C4DM) Queen Mary University of London</institution>
          ,
          <addr-line>London E1 4NS</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper discusses an extension to the Audio E ect Ontology (AUFX-O) for the interdisciplinary classi cation of audio e ect types. The ontology extension implements a uni ed classi cation system that draws on knowledge from di erent music-related disciplines and is designed to facilitate the retrieval of audio e ect information based on low-level and semantic aspects. It extends AUFX-O enabling communication between agents from di erent disciplines within the eld of music creation and production. After brie y discussing the ontology, we show how it can be used to e ciently classify and retrieve e ect types.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In modern music production, composers and producers can choose from a large
number of audio e ect implementations of a variety of e ect types.
Identifying suitable e ects for speci c tasks typically remains a time-consuming manual
process. Moreover, the aspects by which audio e ects are described and
classied can di er depending on the discipline in the eld of music production and
composition. This paper introduces an extension to the Audio E ect Ontology
(AUFX-O) we presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], for the interdisciplinary classi cation of audio
e ect types. AUFX-O is designed to describe audio e ect implementations, i.e.
sound transformation devices, and their use in music production process. While
the core ontology already allows for the annotation of e ect types with commonly
used tags, this ontology extension introduces de ned concepts for e ect
description. It implements a uni ed classi cation system that draws on knowledge from
di erent music-related disciplines and is designed to facilitate the retrieval of
audio e ect information based on low-level and semantic aspects. Its design is
aimed at enabling communication between agents adhering to data models with
conceptualisations informed by and supporting di erent disciplines, both
technical and artistic. After brie y discussing the ontology model, we show how an
e ect type can be described based on low-level and high-level descriptors.1
1 Since AUFX-O and its extension are domain-speci c we refer the reader to [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for a
detailed discussion of audio e ect types and their implementation.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Audio E ect Classi cation</title>
      <p>
        Digital audio e ects are used by composers, performers and sound engineers.
They can be classi ed based on di erent aspects that are relevant for these
individual disciplines. For instance, a software developer would put emphasis on
the signal processing techniques applied in the e ect implementation, whereas a
musician may prefer to describe e ects by their perceptual qualities or by
highlevel semantic descriptors. An example for such a discipline-speci c classi cation
is given in [
        <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
        ], where e ects are categorised based on the perceptual attributes
most a ected by a given sound transformation. The perceptual axes are
loudness, pitch, time, space, and timbre.2 Audio e ects may also be classi ed by the
way its parameters are controlled, for instance by way of manual input, low
frequency oscillators (LFOs), or noise generators. Lower level descriptions include
the signal processing strategies applied in a given e ect type's implementation
and their characteristics, such as the processing domain (i.e. time or frequency
domain)3, and the algorithm or model implemented for their realisation. Figure
1 illustrates layers of discipline-speci c classi cations ranging from low-level to
high-level descriptors. In an e ort to unify these classi cation approaches an
interdisciplinary classi cation model has been proposed in [
        <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
        ], which forms the
basis for the AUFX-O classi cation extension.
      </p>
      <p>Digital implementation
technique</p>
      <p>Processing
Domain</p>
      <p>Applied
processing</p>
      <p>Control
Type</p>
      <p>
        Perceptual
Attribute
low level
features
Semantic
descriptors
high level
features
The main classes and properties of the AUFX-O extension are shown in
Figure 2. The extension incorporates the classi cation systems discussed in Section
2, including the classi cation based on perceptual attributes which we
evaluated by conducting listening tests [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The extension links to AUFX-O via the
:effect_type property that links instances of the aufx:Fx class conceptualising
audio e ects on an abstract level (as physical phenomena) to instances of the
EffectType class. Semantic descriptors and the common name of an e ect type
are linked via annotation properties, while object properties relate EffectType
to PerceptualAttribute, ControlType and AppliedProcessing. The latter
can be further described with object properties linking the Implementation
Technique which in turn is linked to its ProcessingDomain. While the
ontology contains individuals of the PerceptualAttribute and ProcessingDomain
classes, other concepts contain subclass structures for further specialisation.
2 for a detailed discussion of music perception see [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
3 Time domain refers to the variation of a signal's amplitude over time; in the frequency
domain signals are represented by a magnitude and phase as a function of frequency.
xsd:string
      </p>
      <p>:name
:semantic_descriptor
aufx:Fx
:effect_type
:ControlType
:control_type</p>
      <p>:EffectType
:applied_processing
:AppliedProcessing
:implementation_technique
:ImplementationTechnique</p>
      <p>
        We illustrate how e ect types are modelled with the example of the chorus
effect. A chorus e ect arti cially simulates the doubling of voices occurring when
several musicians play the same melody in unison and slight pitch, dynamic,
rhythm and timbre di erences arise because the instruments are not physically
identical, nor are perfectly tuned and synchronised. This widely used e ect
provides some "warmth" to a sound, and can be considered an e ect on timbre.
While its usual implementation involves one or many delay lines, with modulated
length and controlled by a white noise, more realistic sounding implementations
use several slightly pitch-shifted and time-scaled versions of the same sound with
re ned models and mixing them together. Listing 1 shows how this knowledge
is encoded in the ontology. The Chorus class is de ned as a subclass of the
EffectType class, as well as of several anonymous classes re ecting the
descriptors discussed in Section 2. Anonymous classes further describe the processing
applied and the techniques needed for the e ect's implementation. For instance,
the ontology re ects that for the implementation of the chorus e ect there are
three possible processing strategies, AddingDelayedSignal, TimeScaling and
Transposition. Applying transposition, i.e. the changing of pitch of a sound,
can be achieved by three di erent implementation techniques. These include the
PhaseVocoder technique in the FrequencyDomain and the SOLA technique in the
TimeDomain [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions and Future Work</title>
      <p>
        We discussed the ontological representation of an interdisciplinary classi cation
system for audio e ects implemented as an extension to AUFX-O. We showed
how the extension can add value to the ontology by allowing users and agents
to retrieve audio e ect implementations based on di erent characteristics that
are important for di erent disciplines in the eld of music composition and
production. Future work includes extending the ontology with a more thorough
coverage of e ects used in music production, and further development of the
linked data service we presented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], exposing metadata about audio e ect
implementations.
      </p>
      <p>Acknowledgments. This paper has been supported by EPSRC Grant EP/
L019981/1, Fusing Audio and Semantic Technologies for Intelligent Music
Production and Consumption. Mark B. Sandler acknowledges the support of the
Royal Society as a recipient of a Wolfson Research Merit Award.</p>
      <sec id="sec-3-1">
        <title>Class: Chorus</title>
      </sec>
      <sec id="sec-3-2">
        <title>Annotations:</title>
        <p>semantic_descriptor "warm sound",
semantic_descriptor "several performers" ,
name "Chorus"</p>
      </sec>
      <sec id="sec-3-3">
        <title>SubClassOf: EffectType, applied_processing some (AddingDelayedSignal or TimeScaling or Transposition), control_type some WhiteNoiseControl, perceptual_attribute value Timbre</title>
      </sec>
      <sec id="sec-3-4">
        <title>Class: DelayLine SubClassOf:</title>
      </sec>
      <sec id="sec-3-5">
        <title>ImplementationTechnique and (processing_domain value TimeDomain)</title>
      </sec>
      <sec id="sec-3-6">
        <title>Class: PhaseVocoder SubClassOf:</title>
      </sec>
      <sec id="sec-3-7">
        <title>ImplementationTechnique and (processing_domain value FrequencyDomain)</title>
      </sec>
      <sec id="sec-3-8">
        <title>Class: SOLA SubClassOf:</title>
      </sec>
      <sec id="sec-3-9">
        <title>ImplementationTechnique and (processing_domain value TimeDomain)</title>
      </sec>
      <sec id="sec-3-10">
        <title>Class: TimeScaling SubClassOf:</title>
      </sec>
      <sec id="sec-3-11">
        <title>AppliedProcessing and (implementation_technique some (AdditiveModel or PhaseVocoder or SOLA))</title>
      </sec>
      <sec id="sec-3-12">
        <title>Class: Transposition SubClassOf:</title>
      </sec>
      <sec id="sec-3-13">
        <title>AppliedProcessing and (implementation_technique some (AdditiveModel or PhaseVocoder or SOLA))</title>
      </sec>
      <sec id="sec-3-14">
        <title>Class: DelayLine SubClassOf:</title>
      </sec>
      <sec id="sec-3-15">
        <title>ImplementationTechnique and (processing_domain value TimeDomain)</title>
      </sec>
      <sec id="sec-3-16">
        <title>Class: AddingDelayedSignal SubClassOf:</title>
      </sec>
      <sec id="sec-3-17">
        <title>AppliedProcessing and (implementation_technique some DelayLine)</title>
      </sec>
      <sec id="sec-3-18">
        <title>Class: AdditiveModel SubClassOf:</title>
      </sec>
      <sec id="sec-3-19">
        <title>ImplementationTechnique and (processing_domain value FrequencyDomain)</title>
        <p>Listing 1. Excerpt from the AUFX-O classi cation extension relating to the e ect
type Chorus.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Deutsch</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          . (ed.):
          <source>The Psychology of Music, 3rd Edition</source>
          . Academic Press (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Verfaille</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holters</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Z</given-names>
            olzer, U.: Introduction. In: Zolzer, U. (ed.)
            <surname>DAFX - Digital Audio</surname>
          </string-name>
          <string-name>
            <given-names>E ects. 2nd</given-names>
            <surname>Edition</surname>
          </string-name>
          , J. Wiley &amp; Sons (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Verfaille</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guastavino</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Traube</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>An interdisciplinary approach to audio effect classi cation</article-title>
          .
          <source>Proceedings of the 9th International Conference on Digital Audio E ects (DAFx-06)</source>
          , Montreal, Canada (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wilmering</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fazekas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandler</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          :
          <article-title>AUFX-O: Novel methods for the representation of audio processing work ows</article-title>
          .
          <source>Proceedings of the 15th International Semantic Web Conference (ISWC</source>
          <year>2016</year>
          )
          <article-title>(</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wilmering</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fazekas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandler</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          :
          <article-title>Audio e ects data on the semantic web</article-title>
          .
          <source>Presented at the 139th Audio Engineering Society Convention</source>
          , New York, NY, USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Wilmering</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fazekas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandler</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          :
          <article-title>Audio e ect classi cation based on auditory perceptual attributes</article-title>
          .
          <source>presented at the AES 135th Convention</source>
          , New York, NY, USA. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Zolzer, U. (ed.): DAFX - Digital Audio E ects. J. Wiley &amp; Sons, 2nd edn. (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>