<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Harnessing Crowds and Experts for Semantic Annotation of the Qur'an</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amna Basharat</string-name>
          <email>amnabash@uga.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khaled Rasheed</string-name>
          <email>khaled@uga.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>I. Budak Arpinar</string-name>
          <email>budak@uga.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science University of Georgia Athens</institution>
          ,
          <addr-line>GA, 30602</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we illustrate how we harness the power of crowds and specialized experts through automated knowledge acquisition work ows for semantic annotation in specialized and knowledge intensive domains. We undertake the special case of the Arabic script of the Qur'an, a widely studied manuscript, and apply a hybrid methodology of traditional 'crowdsourcing' augmented with 'expertsourcing' for semantically annotating its verses. We demonstrate that our proposed hybrid method presents a promising approach for achieving reliable annotations in an e cient and scalable manner, especially in cases where a high level of accuracy is required in knowledge intense and sensitive domains.</p>
      </abstract>
      <kwd-group>
        <kwd>semantic annotation</kwd>
        <kwd>disambiguation</kwd>
        <kwd>classi cation</kwd>
        <kwd>ontology</kwd>
        <kwd>Qur'an</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Thematic annotation of religious texts, in particular, the classical sources of
knowledge in the Islamic domain in the Arabic language, has not received much
attention, partly owing to time and knowledge constraints from experts required
for such an annotation process. In our research, we consider the application of
specialized human computation methods such as nichesourcing ([
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) in an
an attempt to scale this process of annotation. Nichesourcing or
expertsourcing extends the idea of engaging skilled and knowledgeable persons in place of
faceless crowds for human driven tasks. We employ nichesourcing as means of
augmenting traditional crowdsourcing methods rather than as an alternate.
      </p>
      <p>
        In this paper, we show the results of an exploratory study focussing on
two knowledge intensive tasks: the thematic disambiguation and annotation of
Qur'anic verses using its Arabic script. While this has been tackled through
a pure crowdsourcing approach in our earlier work in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], we determined that
several tasks are rather knowledge intensive and require domain expertise. In
this case, not only the knowledge of Qur'anic Arabic is considered imperative,
the annotation of the Arabic verses also requires understanding the context and
content of the given verse.
Hybrid Architecture for Harnessing Crowd and Expert
Annotations
We design and develop a hybrid work ow architecture that connects a
crowdsourcing framework with an expertsourcing application as shown in Figure 1.
Crowdsourcing Stage: We design a task management engine that is
responsible for generating tasks, retrieving and aggregating results. The tasks are
published on the Amazon Mechanical Turk (AMT)1 platform. A complete
workow management system is implemented (a derivative of a work ow model for
Linked Data Management presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]), which includes means for generating
dynamic tasks from a range of task pro les. The semantic annotation process is
driven by an ontology schema. The task input is generated by retrieving relevant
candidate verses from the available external data sources such as the
SemanticQuran [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] dataset.
      </p>
      <p>The AMT crowd performs the thematic disambiguation and thematic
annotation tasks. Both tasks are based on the Arabic script of the Qur'an. For the
disambiguation task, a question is presented to the crowd, which includes a verse,
along with a highlighted, candidate explicit assertion for the given theme, and
the crowd responds by declaring this assertion as either a positive or negative by
determining if the occurrence is a true occurrence of the given theme. The
annotation tasks require deeper knowledge and understanding of the Arabic text.
The crowd determines whether the given verse contains any implicit reference
to the given theme. If their response is positive, then they are also required to
provide the portion of the verse (a meaningful phrase or a word) that implies
the presence of the theme. As a form of a quality measure, the crowd is also
1 http://www.mturk.com</p>
      <p>Harnessing Crowds and Experts for Semantic Annotation of the Qur'an
required to provide a con dence level (ranging from Very High to Very low), to
indicate their con dence in their response.</p>
      <p>Decision Analytics: We collect and aggregate the responses based on
statistical measures of aggregation. Weighted con dence measures and thresholds are
applied. Based on this aggregation, the completed tasks are marked as either
Approved or Reviewable. A high con dence and aggregation threshold is applied
for the approved tasks. This decision analytics results in identifying the
candidate tasks for expertsourcing. The tasks marked as reviewable, which fail to
meet the agreement thresholds, are sent o for expert annotations.</p>
      <p>Expertsourcing Stage: For this purpose we designed a custom web application
to engage with experts. A RestAPI connects the crowdsourcing task management
engine with the expertsourcing application. The tasks are sent to the remote
application and experts are noti ed when the tasks become available. The experts
also see the candidate responses collected from the crowdsourcing stage. The
experts have either the option to choose from the available annotations (collected
during the crowdsourcing stage) or provide their own if they do not agree with
either one. An example is shown in Fig 2. We present the same task to three
experts to analyze the annotation agreements. The approved and validated
annotations are passed on for Ontology Population and linked with existing data
sources.
The experimental setup assigned each task to 5 crowd workers. For the reviewable
crowd tasks that were sent to experts, 3 experts were assigned to each reviewable
task. Table 1 shows the results obtained.</p>
      <p>Crowd Tasks Expert Tasks
Task Disambiguation Annotation Disambiguation Annotation
Approved 1267 477 34 96
Reviewable 40 107 6 11
Total 1307 584 40 107</p>
      <p>Table 1. Results for Disambiguation and Annotation Tasks</p>
      <p>The results of our exploratory study provide interesting insights into the
application of human computation methods to knowledge intensive tasks. Our
task design involved the thematic disambiguation and annotation of the Qur'anic
verses based on the original Arabic script. For the disambiguation task, 99% of
the tasks were able to reach an agreement by combining contributions of crowds
and experts. Only about 3% tasks needed expert contributions. For the
annotation tasks, about 18% tasks needed expert contributions. There were 10% tasks
that did not reach an agreement with both crowd and expert contributions. An
administrative review of these cases indicate that some annotations are a matter
of personal taste and judgement and closed agreement is therefore di cult. Most
of these annotations cannot be classi ed as wrong, nor better than the others
based on an automated agreement mechanism.</p>
      <p>Our knowledge acquisition and review work ow to selectively elicit expert
annotations only where needed indeed presents a promising method. We utilize
annotation agreement and distance analytics to route the appropriate tasks that
need expert contributions. Our results suggest that such a hybrid approach
indeed creates for a more accurate and reliable annotation process. This method
can be e ectively utilized for qualitative dataset management and semantic
annotation tasks in an economic and feasible manner through crowd engagement,
while reducing the need for expert contributions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>De Boer</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hildebrand</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aroyo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Leenheer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dijkshoorn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tesfa</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreiber</surname>
          </string-name>
          , G.:
          <article-title>Nichesourcing: Harnessing the power of crowds of experts</article-title>
          .
          <source>In: Knowledge Engineering and Knowledge Management</source>
          . Springer (
          <year>2012</year>
          )
          <volume>16</volume>
          {
          <fpage>20</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Oosterman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bozzon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Houben</surname>
            ,
            <given-names>G.J.e.a.</given-names>
          </string-name>
          :
          <article-title>Crowd vs. experts: nichesourcing for knowledge intensive tasks in cultural heritage</article-title>
          ,
          <source>Int. WWW Conferences Steering Committee</source>
          (
          <year>2014</year>
          )
          <volume>567</volume>
          {
          <fpage>568</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Basharat</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arpinar</surname>
            ,
            <given-names>I.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rasheed</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Leveraging crowdsourcing for the thematic annotation of the qur'an</article-title>
          .
          <source>In: Proceedings of the 25th International Conference Companion on World Wide Web, International World Wide Web Conferences Steering Committee</source>
          (
          <year>2016</year>
          )
          <volume>13</volume>
          {
          <fpage>14</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Basharat</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arpinar</surname>
            ,
            <given-names>I.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dastgheib</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kursuncu</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochut</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dogdu</surname>
          </string-name>
          , E.:
          <article-title>Semantically enriched task and work ow automation in crowdsourcing for linked data management</article-title>
          .
          <source>International Journal of Semantic Computing</source>
          <volume>8</volume>
          (
          <issue>04</issue>
          ) (
          <year>2014</year>
          )
          <volume>415</volume>
          {
          <fpage>439</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Sherif</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.C.N.</given-names>
          </string-name>
          :
          <article-title>Semantic Quran - a multilingual resource for naturallanguage processing</article-title>
          .
          <source>Semantic Web</source>
          <volume>6</volume>
          (
          <issue>4</issue>
          ) (
          <year>2015</year>
          )
          <volume>339</volume>
          {
          <fpage>345</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>