<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>human-in-the-loop: scaling occupation taxonomy at Indeed</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Suyi Tu</string-name>
          <email>suyitu@indeed.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olivia Cannon</string-name>
          <email>ocannon@indeed.com</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>A hierarchical occupation taxonomy helps Indeed better match job seekers and jobs. Historically, scaling this hierarchical taxonomy in international markets was a time intensive and highly manual process. Leveraging the strengths of both machine learning models and subject matter experts led to the creation of an improved human-in-the-loop system that met the needs of a growing business without sacrificing quality. This paper describes this system and discusses the challenges and insights from its implementation. Specifically, it highlights the value of involving subject matter expertise beyond just labeling and therefore ofers the application of an expert-in-the-loop framework for scaling taxonomy. Human-in-the-loop, Expert-in-the-loop, taxonomy, subject matter expert, natural language processing, scaling matching between jobs and job seekers and enables per- Figure 1: Example of an occupation taxonomy branch omy is granular, precise, and customizable by market. leverages broader and narrower concept relationships to RecSys in HR'22: The 2nd Workshop on Recommender Systems for</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>As the #1 job site in the world, Indeed hosts a massive</title>
        <p>volume of unstructured data in the form of millions of
job descriptions and resumes. Classifying a job or work
experience as an occupation is one way of extracting
structured data from these documents. This improves
sonalization.</p>
        <p>The Indeed Taxonomy team has spent years
thoroughly researching and developing a hierarchical
occupation taxonomy and a hand-curated rule system to classify
jobs for core strategic markets. The occupation
taxon</p>
      </sec>
      <sec id="sec-1-2">
        <title>Consequently, it required significant time and resource</title>
        <p>investment for international scaling and ongoing
maintenance. As demand for Indeed’s presence to expand
into more international markets rapidly grew, the team
needed new ways to scale this work faster, with fewer
resources, and without sacrificing data quality.</p>
        <p>This paper reflects on the challenges and successes
encountered while introducing a human-in-the-loop (HITL)
approach to scaling Indeed’s occupation taxonomy
development and classification to support its rapidly
expanding international business strategy. In particular, it
discusses a core learning: that targeted combination of
domain expertise and broad application of machine learning
technology has been the key to success. The scope of this
combination exceeds the level of human-augmentation
generally implied by HITL, and we therefore describe it
more precisely as expert-in-the-loop (EITL).
Human Resources, in conjunction with the 16th ACM Conference on
[Hierarchy visualization of part of food &amp; beverage
occupation taxonomy]
2.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Motivation</title>
      <sec id="sec-2-1">
        <title>The occupation taxonomy at Indeed is a hierarchy that</title>
        <p>enrich understanding of occupation types. For example,
Bartenders is narrower than its broader taxonomy
concept, Food &amp; Beverage Servers, and Food &amp; Beverage Servers
is narrower than its broader concept, Food &amp; Beverage
Occupations. It is therefore understood that Bartenders is a
type of Food &amp; Beverage Occupation. This broadest
grouping is referred to as a Top Level concept. Given a job or
resume document, our classification system assigns the
most specific occupation concepts possible.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Indeed is a global company with a presence in over</title>
        <p>60 countries. Occupation taxonomy design in diferent
markets may vary due to factors such as employment
landscape, economy, and the volume of jobs hosted on</p>
      </sec>
      <sec id="sec-2-3">
        <title>Indeed in that location. However, the Indeed Taxonomy</title>
        <p>team has found that occupation taxonomy design is
typically more similar than it is diferent across markets.</p>
        <p>Compared with occupation taxonomy concepts that exist
in the US, the most similar International market with
complete taxonomy coverage has over 90% of occupation
concepts in common, while even the least similar
International market has roughly 50% in common. Having a
monality and distinction allows for sharing knowledge
nEvelop-O
among markets, designing market-specific features, and
avoiding excessive granularity where it is not relevant
or required.</p>
        <p>The Taxonomy team at Indeed is led by taxonomists
who are highly skilled professional researchers and
possess graduate degrees in Information Science, Philosophy,
Linguistics, or related disciplines. Teams of Taxonomy
Analysts are hired specifically to support strategically
significant markets, and these individuals typically possess
work experience in data analysis or professional research.</p>
        <p>They must have strong expertise in a market’s
employment landscape as well as spoken and written fluency
in the oficial languages. They either possess a formal
education in Information Science with hands-on
experience with Taxonomy, or receive extensive on-the-job
training in Information Science concepts and Taxonomy [End-to-end workflow of the expert-in-the-loop system]
best practices. Their work is informed by consulting Figure 2: High-level overview of the expert-in-the-loop
sysrelevant quantitative and qualitative data, competitive tem. Circled steps are heavily SME involved; Squared steps
analyses, and applying many of the same user-centered are fully automated.
design principles seen in UX Research or Content Design.</p>
        <p>Indeed’s occupation taxonomy expansion for the first
few international markets was based on a rule-populated can refer to diferent model architectures in the various
manual approach and took years to achieve minimal vi- development phases and markets. This section discusses
able state in each new market. As international business model selection and evolution in detail.
exponentially expanded, the human-intensive process of When entering a new market, instead of manually
crecrafting separate rule systems for each new market and ating a rule system from scratch for occupation
classificakeeping them up-to-date became insurmountable. tions, we employ machine learning. We start by selecting</p>
        <p>
          There is plenty of research on applications of machine a model for training that makes the best use of the
exlearning techniques to large-scale taxonomy develop- isting data. This existing data may include SME-curated
ment and document classification in diverse fields with occupation hierarchies and high-quality, rule-generated
minimal human intervention [
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1, 2, 3, 4</xref>
          ]. However, in labels in existing markets. Linguistic and cultural
resemorder to enable a consistent and high-quality user experi- blance between markets are two main factors guiding
ence around the globe, it is important to identify shared the model selection decision at this initial phase, because
occupations between markets and develop new compo- labels of the same language are often good training data
nents of the taxonomy to reflect market-specific occu- sources and cultural resemblance is often reflected in the
pation types or relationships, which involves research overlapping occupation concepts.
from domain experts. In addition, the system in produc- There are several model options that each come with
tion requires close monitoring to ensure data quality and their own benefits and tradeofs at diferent phases.
freshness [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. These challenges made adopting a purely Model options include a multilingual BERT (M-BERT)
automated system less desirable. There is also emerging model [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], and convolutional neural networks (CNN)
research on the value of incorporating domain expertise [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ] with diferent training data. For new markets
in facilitating efective automation in highly technical that share a language and concepts with established
marifelds [
          <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
          ]. We propose that occupation taxonomy kets, a CNN model is trained to make initial predictions.
development and design similarly benefits from this ex- If there is no resemblance in either aspect, we default to
pert augmented approach that leverages the strength of a global M-BERT model that is trained with high quality
both machine learning models and subject matter experts data from all the existing markets, regardless of linguistic
(SMEs). In the following sections, “experts” in our EITL or cultural similarity. Since the generalization power of
system are referred to as “SMEs”. M-BERT comes with tradeofs of large model size and
high inference latency, we have to limit the input length
3. Expert-in-the-loop Taxonomy and therefore limit the prediction accuracy.
        </p>
        <p>After the selected model generates predictions, SMEs
Scaling conduct research and build out the market-specific
occupation taxonomy iteratively. For example, SMEs identify
novel market-specific occupations like Judicial Scriveners
We represent a simplified overview of our EITL system
in Figure 2. Note that the “model” specified in the figure
in Japan, Pizza Chefs in Italy, Hostel Wardens in India, or 4.1. Transfer subject matter expertise to
removing occupations that do not apply to the specific understanding in data
market. The ongoing research is reflected by a dynamic
layer at the inference time, which is used to process pre- Shadowing taxonomy SMEs and listening to their
indictions into the market-specific taxonomy. This step is sights and challenges ofered valuable information that
referred to as ”Predictions fallback” in Figure 2. would not otherwise have been obvious. Incorporating</p>
        <p>To determine deployment eligibility, Taxonomy SMEs this information into the modeling process resulted in
label the model predictions for evaluation on a concept- significant impact.
by-concept basis and push concepts to production based For example, we learned that although taxonomy
hion a predetermined quality threshold. Since the top level erarchies may vary in structure from market-to-market,
taxonomy concepts are typically shared among all mar- there are occupation concepts across these markets that
kets, it takes a short amount of time for a new market to are conceptually identical but have diferent IDs due to
deploy top level concept predictions in production. those structural discrepancies. This posed an issue for</p>
        <p>For novel, market-specific concepts, or for concepts cross-market training and prediction. By incorporating
where prediction quality is below threshold, SMEs pro- the mapping into model training, we saw gains in both
vide training labels for retraining the model. For de- precision and recall in all the locales, with an average
ployed models and concepts, we follow a monitoring increase of 5% in precision and 12% in recall.
process to discover and act on issues. Learning about diferences in taxonomy concepts</p>
        <p>After a market-specific taxonomy is built out, our re- across markets also prompted us to explore and apply
search shows that a market-specific CNN model utilizing market-specific models after the initial set of
marketthis knowledge with translated training data often outper- specific hierarchies are built out. Experiments in the
forms the global M-BERT model. Therefore, the “model” initial adopted market showed an average increase of
in Figure 2 will be replaced by a market-specific model 12% in precision when the threshold is set to keep recall
for subsequent loops. the same.</p>
        <p>This EITL approach leveraging machine learning
models eliminates the most expensive task in the process: 4.2. Quality training and evaluation data
the need to curate and maintain a rule-based system in at scale
pursuit of capturing 100% of all possible classifiable
syntaxes. It allows the Indeed Taxonomy team to divert Quality data is key for training and evaluating any
auresources to more SME-critical tasks, such as research- tomated systems. With limited resources, we explored
ing and creating market-specific occupation hierarchies outsourcing labeling and using user feedback data to
and SME-annotated datasets, expediting deployment, and directly approach the problem, and also revisited
prioriother metadata projects. We also benefit greatly by in- tization of the tasks based on our learning.
volving SME knowledge beyond just labeling like most
traditional HITL systems, which we’ll discuss in detail 4.2.1. Outsourcing labeling
in the next section. As a result of this approach, we
observed a six times faster deployment with similar SME
resources was achieved in markets adopting this system
(compared to the fully manual process), and this system
is now in production in dozens of markets.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Crowd-sourcing labels for tasks where SME knowledge</title>
        <p>
          is needed has been known to be challenging [
          <xref ref-type="bibr" rid="ref1 ref12 ref13">1, 12, 13</xref>
          ].
        </p>
        <p>After working with external labelers on multiple tasks
with mixed outcomes in the past years, we learned that
significant work needed to be put in upfront in order to
gain high quality outcomes–underscoring the diference
4. Challenges and Insights between EITL and HITL. Certain tasks need strong
domain specific subject matter expertise which is hard to
We learned a lot from implementing this approach. In gain with short-term training sessions, particularly when
this section, we will discuss our success in leveraging much of the communication is asynchronous. Those
subject matter expertise in the modeling process, adding tasks should not be outsourced in the first place.
Guidtransparency to a blackbox system, and establishing a ing labelers to evaluate predictions required ongoing
monitoring process. We will also share insights from ex- support in applying consistent classification heuristic to
ploring diferent data sources and identifying the unequal their labeling tasks. For example, coaching labelers to
nature of diferent errors. understand the diference between occupation and work
environment was a unique challenge. Frequently, jobs
for School Custodians would be labeled as a type of
Education &amp; Instruction Occupation due to their workplace
environment. Occupationally, however, this should be a
Cleaning &amp; Grounds Maintenance Occupation. Moreover, our close collaborations with SMEs on labeling tasks, we
if a task involves distinguishing between specific licenses learned that there is usually a certain level of subjectivity
or specialties in diferent industries, significant research associated with decisions, which naturally brought us
and domain knowledge are needed in order to produce this question: How bad are the errors? For example, if the
high quality labels. model makes the same error as a non-SME user, it likely</p>
        <p>After identifying tasks suitable to outsource, the next does not hurt user experience as much as a completely
steps are to work with SMEs to scope the task, provide out-of-place error. After examining diferent definitions,
clear guidelines, and set up monitoring-QA-calibration considering trade-ofs between information granularity
loops to help ramp up the quality. In a labeling task and interpretability, we ended up defining three levels
where the goal is to generate evaluation data for top of error severity with easy interpretations: 1) the mild
level occupation predictions in a new market, we found errors are within 2 hops of the correct labels, typical
it to be more eficient to scope the task in the binary examples include siblings, parents and children nodes;
fashion: During the labeling task, labelers are asked to 2) the moderate errors are at the same top levels with the
label the predictions such as “Job A is a type of Education correct labels, usually due to granularity issues; and 3)
Instruction Occupations job” as correct or incorrect. The the severe errors are in diferent top levels, which will
incorrect predictions are then sent to SMEs to assign result in bad user experience if not fixed.
correct labels. This second step is essential to improve Having this insight allows the SME team to conduct
model quality but extremely hard for non-SMEs since more nuanced error analysis such that they can better
it requires familiarity with definitions of a few dozen prioritize severe errors for quality improvement tasks. It
occupation concepts. also allows the data science team to measure quality in
a more practical manner, and utilize data from multiple
4.2.2. User input data sources which would be considered too noisy otherwise.</p>
      </sec>
      <sec id="sec-2-5">
        <title>User input data comes in large volume, but tends to be</title>
        <p>
          noisy [
          <xref ref-type="bibr" rid="ref14 ref4">4, 14</xref>
          ]. After reviewing user input data against 4.4. Establish process of monitoring,
SME labels, we learned that there could be discrepancies diagnosis, and action on issues
associated with understandings of job descriptions and With machine learning models in production supporting
concept definitions, often resulting in diferent or miss- dozens of locales and thousands of taxonomy concepts in
ing labels. In a model-assisted user input task, we found total, we needed to monitor performance of the models
20-30% of user overrides disagreed with SME labels. In and act on issues in a timely manner. Therefore we are
addition, user input was easily influenced by UI design. exploring a workflow with SMEs, which involves a
monA recent UI change making the editing option less ob- itoring dashboard with alerts on the individual concept
vious resulted in a 4% decrease in users overriding the level and a process to diagnose and act on issues. The
predictions. These challenges reflect the quality tradeof alerts are designed to be triggered when a drastic drop
associated with this free data source. in performance is seen in a node with suficient labels.
Those alerts will then be sent to SMEs in each locale,
4.2.3. Prioritization of labeling tasks marked with an urgency level based on the scale of the
We learned that, in terms of labeling, quality usually drop. The SMEs will follow a triaging decision tree,
leadcomes with the tradeof of cost and speed, especially ing to either ignoring the alert or creating a ticket with
when the dimension of label space is in the thousands. their observations and judgments for data scientists or
With that learning, we decided to prioritize SME re- engineers to resolve.
sources for evaluation datasets, while using alternative
data sources for training, such as taxonomy data from 4.5. Strive for transparency into
other markets and user input. When certain classes black-box systems
proved to perform poorly with the alternative training
data source, we then prioritized labeling better training
data for those specific classes. With this decision, we
were able to significantly decrease time to the first model
deployment and evaluation.
        </p>
      </sec>
      <sec id="sec-2-6">
        <title>Deep learning systems are often treated as black boxes,</title>
        <p>but this technical opacity can lead to ineficient
crossfunctional collaboration. For example, during the initial
collaboration period, SMEs were tasked to provide labels
for a model in a market with around a thousand concepts.</p>
        <p>After a significant amount of work was put in, we realized
4.3. Understand what error matters the evaluation metrics were inaccurate due to nonrandom
In the classic machine learning model measurements, selection of labeled documents. In addition, the initial
error is defined to be a binary concept. However, with instruction of providing a large fixed amount of labels per
concept could be improved since concepts that perform
[Example of the token importance function applied to a job in</p>
        <p>the internal tool]
well do not need additional training data, and the ones
with fewer jobs struggle to find enough documents to be
labeled and have little efect on the overall performance.</p>
        <p>
          We started with knowledge sharing, Q&amp;A, and
discussion sessions to connect common operation actions
with implications on the models, and gradually built out
resources and documentation for FAQs. In addition, we
built a function to visualize token importances in a given
training sample with the Integrated Gradients [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]
technique, which is used by SMEs to promote transparency
into why a document might be predicted as a given
occupation.
        </p>
        <p>These insights and successes could not have been
achieved without close collaboration and
communication. It enabled us to efectively distribute SME resources
to the highest priority task, which resulted in the
research, development, and publishing of over 700 existing
or new occupation taxonomy concepts in strategically
significant new markets within six months.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Conclusion and future work</title>
      <p>In this paper, we presented a human-in-the-loop system
at Indeed that scales occupation taxonomy development
and application to international markets with high
quality outcomes. Over the course of developing the system,
we found tremendous value in leveraging subject matter
expertise throughout the process, ranging from data
collection and processing, model training and evaluation,
to performance monitoring and diagnosis. Therefore we
propose calling this system expert-in-the-loop.</p>
      <p>In ongoing work, we are looking to partner more
closely among functions to allow for combined SMEs,
explore further improvements to the current machine
learning models, and build better functionality for
overriding training and production data.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>We would like to thank everyone on the Metadata &amp;</title>
        <p>Taxonomy team that contributed to this system for their
diligent work and thoughtful collaboration. We are
grateful to the reviewers of this paper for providing valuable
feedback and comments.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Strube</surname>
          </string-name>
          ,
          <article-title>Deriving a large-scale taxonomy from wikipedia</article-title>
          ,
          <source>in: AAAI</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Probase: a probabilistic taxonomy for text understanding</article-title>
          ,
          <source>Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Javed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>McNair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jacob</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <article-title>Carotene: A job title classification system for the online recruitment domain</article-title>
          ,
          <source>2015 IEEE First International Conference on Big Data Computing Service and Applications</source>
          (
          <year>2015</year>
          )
          <fpage>286</fpage>
          -
          <lpage>293</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Das</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Levine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. D.</given-names>
            <surname>Fabbrizio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Datta</surname>
          </string-name>
          ,
          <article-title>Large-scale taxonomy categorization for noisy product listings</article-title>
          ,
          <source>2016 IEEE International Conference on Big Data (Big Data)</source>
          (
          <year>2016</year>
          )
          <fpage>3885</fpage>
          -
          <lpage>3894</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Noy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Patterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          ,
          <article-title>Industry-scale knowledge graphs: Lessons and challenges: Five diverse technology companies show how it's done</article-title>
          ,
          <source>Queue</source>
          <volume>17</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Gennatas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Ungar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pirracchio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Eaton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Reichmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Interian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Luna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. B.</given-names>
            <surname>Simone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Auerbach</surname>
          </string-name>
          , E. Delgado,
          <string-name>
            <given-names>M. J. van der</given-names>
            <surname>Laan</surname>
          </string-name>
          , T. D.
          <string-name>
            <surname>Solberg</surname>
            ,
            <given-names>G. Valdes,</given-names>
          </string-name>
          <article-title>Expertaugmented machine learning</article-title>
          ,
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>117</volume>
          (
          <year>2020</year>
          )
          <fpage>4571</fpage>
          -
          <lpage>4577</lpage>
          .
          <source>doi:1 0 . 1 0</source>
          <volume>7 3</volume>
          / p n a
          <source>s . 1 9</source>
          <volume>0 6 8 3 1 1 1 7 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Wallace</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Small</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Brodley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Trikalinos</surname>
          </string-name>
          ,
          <article-title>Deploying an interactive machine learning system in an evidence-based practice center: Abstrackr</article-title>
          , in
          <source>: Proceedings of the 2nd ACM SIGHIT International Health Informatics Symposium</source>
          , IHI '12,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2012</year>
          , p.
          <fpage>819</fpage>
          -
          <lpage>824</lpage>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 2 1 1 0 3 6 3 . 2 1 1 0 4 6 4 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Interactive learning for eficiently detecting errors in insurance claims</article-title>
          ,
          <source>in: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , KDD '11,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2011</year>
          , p.
          <fpage>325</fpage>
          -
          <lpage>333</lpage>
          . doi:
          <volume>10</volume>
          .1145/2020408.2020463.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          , CoRR abs/
          <year>1810</year>
          .04805 (
          <year>2018</year>
          ). URL: http://arxiv.org/abs/
          <year>1810</year>
          .04805. arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>LeCun</surname>
          </string-name>
          , Y. Bengio,
          <article-title>Convolutional networks for images, speech</article-title>
          , and time series,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Convolutional neural networks for sentence classification</article-title>
          ,
          <source>in: EMNLP</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rampalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doan</surname>
          </string-name>
          , Chimera:
          <article-title>Large-scale classification using machine learning, rules, and crowdsourcing</article-title>
          ,
          <source>Proc. VLDB Endow</source>
          .
          <volume>7</volume>
          (
          <year>2014</year>
          )
          <fpage>1529</fpage>
          -
          <lpage>1540</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R.</given-names>
            <surname>Bekkerman</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Gavish, High-precision phrasebased document classification on a modern scale</article-title>
          ,
          <source>in: KDD</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-D.</given-names>
            <surname>Ruvini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Sarwar</surname>
          </string-name>
          ,
          <article-title>Large-scale item categorization for e-commerce</article-title>
          ,
          <source>Proceedings of the 21st ACM international conference on Information and knowledge management</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sundararajan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Taly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Axiomatic attribution for deep networks</article-title>
          ,
          <source>ArXiv abs/1703</source>
          .01365 (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>