<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Taking stock: a Linked Data inventory of Compliance Checking terms derived from Building Regulations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ruben Kruiper</string-name>
          <email>r.kruiper@northumbria.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IoannisKonsta</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alasdair J.GG.ray</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>FarhadSadeghineko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>RichardWatson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>BimalKumar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>AutomatedComplianceChecking,KnowledgeGraph, NaturalLanguageProcessing,LinkedData</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Architecture and Built Environment Northumbria University</institution>
          ,
          <addr-line>Newcastle</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Mathematics and Computer Sciences Heriot-Watt University</institution>
          ,
          <addr-line>Edinburgh</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <fpage>222</fpage>
      <lpage>237</lpage>
      <abstract>
        <p>ComplianceChecking (CC) would be a loteasier if we couldautomaticallmyap between(1) terms thatoccur in buildingregulationsand(2) elementsof buildingsandbuildingproducts.However, the terminologyusedintheregulationiss vastlydiferent from the terminology found in Building Information Models(BIM). We are thereforeforcedtosomehow shoehorn thevocabularyof regulatorytermsintoa setof classesthatmay wellbe severalorders of magnitudesmaller.This paper aimstoreducethegap betweentermsfound verbatimin theregulationsa,ndtheclassesthatexistin LinkedDataVocabularies in ArchitectureandConstructionW.e exploretheautomatedextractionof domainterminologyfrom buildingregulationsa,ndinterlinktheresultingtermswithexistingcontrolledvocabularieslikeUniclass. The resultingKnowledgeGraph (KG) canbe used tosuggestrelevantandrelateddomainterminology, which improves collectintgheinventoryof LinkedDatatermsrequiredfor CC.</p>
      </abstract>
      <kwd-group>
        <kwd>Building Regulations</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Most building work requires approval, hence regulations are frequently accessed by professionals
in the construction industry – from architects to building inspectors and cont1r]a.cOtnors [
top of this, regulatory compliance is of critical importance to existing building, as corroborated
by incidents like the disastrous Grenfell 2fire]. [Despite a plethora of motivatio3n,s4[
        <xref ref-type="bibr" rid="ref5">, 5</xref>
        ], and
a long history of research6,[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a solution to Automated Compliance Checking (ACC) remains
at large – also see our other paper submitted to this works8h].opTh[e crux of most ACC
motivations is to improve the usability of the building regulations – in terms of efectiveness,
eficiency and ease-of-use. A related strand of ACC research specifically focuses on supporting
human experts during a compliance audit or even during des9i,g1n0[
        <xref ref-type="bibr" rid="ref11">, 11</xref>
        ]. We believe that
providing support during Compliance Checking (CC) is a viable approach, in contrast to ACC.
      </p>
      <p>In this paper we explore the development of a domain lexicon for CC, captured in a Linked
Data Knowledge Graph (KG). Such a lexicon could greatly benefit research on intelligent
∗Correspondingauthor.</p>
      <p>CEUR
CEUR
Workshop
Proceedings</p>
      <p>ceur-ws.org
ISSN1613-0073
Regulatory Compliance (i-ReC) tools, e.g., by easing the mapping between the regulatory texts
and CC rules. However, manually constructing and maintaining even small parts of this lexicon
is a tedious task, see AppendAix. Our aim is to speed up the identification of a comprehensive
set of relevant concepts, related terms and surface forms – the form in which a concept occurs in
text. Sectio2ndescribes related work on deriving a lexicon for CC from text. Sec3tdieosncribes
our aim, the data we use, and our approach to suggesting related terms to a human annotator.
Section4 provides some insight in our preliminary results. Despite our relatively small set of
input regulations, we are able to quickly provide annotators with a relatively comprehensive
overview of related terminology and surface forms. We believe our exploratory approach is of
interest to the wider community, and share our data and code so that others may extend and
improve the work as they see fit1.</p>
    </sec>
    <sec id="sec-2">
      <title>2. From domain vocabulary to conceptualisation</title>
      <p>
        In this paper, we useb‘uilding regulations’ to refer to any regulations captured in building
standards, codes of practices, guidance documents and so on. We use the tleexrimcosn‘’ and
‘conceptualisation’ interchangeably to refer to the set of classes and properties required to
compose compliance checking rules. Deriving a conceptualisation from text is part of any
approach to ontology12[] or taxonomy 1[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] learning from text. Expected steps include term
extraction, concept identification, and internal and external concept linking.
      </p>
      <p>Term Extraction: This step revolves around identifying which words, or groups of words,
denote a term – a single unit of information. Related Natural Language Processing tasks
include Information Extraction (I1E4)][, Semantic Role Labelling (SRL1)5[], and domain term
discovery [16]. In the ACC domain relevant studies include tagging of regulations with semantic
markup, either manually, e.g., RASE17[], or automatic, e.g., Semantic Information Elements
[18]. One challenge identified by these studies is handling terms that consist of multiple words,
so-called Multi-Word Expressionss (MWE1)9][. Proper handling of MWEs is a key issue in NLP
[20, 21, 22], and is especially relevant for IE in technical domains [23].</p>
      <p>
        Concept Identification: This step involves turning the most salient terms into the classes
and properties of a conceptualisation. Related NLP tasks include entity res2o4l,u2t5i]oann[d
canonicalisation26[]. Ideally, concepts are provided with a domain-specific definition and a
set of surface forms 1[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Beyond existing ontologies and classification schemes, such as the
Building Topology Ontology (BOT2)7[] and Uniclass [
        <xref ref-type="bibr" rid="ref14">28</xref>
        ], useful resources include domain
dictionaries. While construction domain dictionaries exist, even as part of the British Standards,
their terms and definitions are often proprietary – limiting their use in generating a shared
conceptualisation for CC. To the best of our knowledge there exists no work on automated
concept identification in the CC domain.
      </p>
      <p>
        Internal Concept Linking: This step refers to the identification of relations between
the concepts in a controlled vocabulary, e.g., a conceptualisation. NLP research exists on
automatically learning hierarchical relations between term2s9,]e..Sgy.,n[onym identification
often relies on similarity measures, e.g.,3[
        <xref ref-type="bibr" rid="ref17 ref18">0, 31, 32</xref>
        ], while hyper/hyponymy (super/subclass)
1We encourage readers to replicate our results and extend our approach, by sharing our code and data at:
https://github.com/rubenkruiper/irec
detection may also rely on syntactic patterns3,3e].g..I,n[ the CC domain WordNet34[] has
been used to link concepts [
        <xref ref-type="bibr" rid="ref21">35</xref>
        ] and work exists on a feature-based relation classifier [
        <xref ref-type="bibr" rid="ref22">36</xref>
        ].
      </p>
      <p>External Concept Linking: This step refers to identifying relations between the concepts
in a controlled vocabulary, and concepts found in other resources like BOT and Uniclass. A
relation of interest is the indication that two classes are identical, e.g., ‘skos:e’xacntdMatch
‘owl:equivalentClass’. Such identity relations make it possible for independently constructed
datasets to use each others’ informati3o7n]. [To the best of our knowledge there exists no work
on automating external concept linking or alignment for CC.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <p>3.1. Aim
AppendixA describes our initial work towards manually developing a conceptualisation for
the CC domain. Three domain experts each identify a group of terms related to a subtopic,
e.g., ‘thermal insulation’. They identify links between terms found in existing vocabularies, and
extend the terminology where they believe this is required. The three annotators note various
issues, including:
• not being sure if terms added to the KG actually occur in the regulations
• not knowing when the collected terms comprehensively describe a small subdomain
• the tediousness of identifying new terms and relations, especially when definitions are
missing and sources may not be reliable
Over a combined 6 days the annotators identify merely 302 vocabulary terms and link them
to 214 external resources. Ouarim is to speed up such manual eforts to identify relevant CC
concepts, related terms and surface forms. It is important to distinguish (1) a vocabulary of
words that occur verbatim in the regulations from a (2) a controlled vocabulary. The former
may simply refer to the set of all the words that occur in the regulations, in our case this set
also includes MWEs. The latter is a set of predefined and preferred terms that may be used in a
domain taxonomy or other knowledge organisation system – the conceptualisation.</p>
      <sec id="sec-3-1">
        <title>3.2. Data and approach</title>
        <p>
          In order to make our code and data available we rely on the UK Merged Approved Documents
[
          <xref ref-type="bibr" rid="ref24">38</xref>
          ], which we will refer to as MAD. This is a set of open access British Building regulations and
guidance that make up a total of 1,247 pages. To help identify which terms are specific to the
building domain, we rely on a background corpus of 5 openly available EU regulations on the
design of medical devices39[
          <xref ref-type="bibr" rid="ref26 ref27 ref28 ref29">, 40, 41, 42, 43</xref>
          ]. As suggested by 4[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] we selected the background
corpus based on the similar text genre and similar time periods of publication1. pTraobvliedes
some insight in the sizes of both input text corpora.
        </p>
        <p>
          Figure 1 illustrates our approach. First, we extract two sets of salient noun-based spans of
words – one from MAD and one from a background corpus of EU regulations on the design of
medical devices. Sectio3n.3 explains how we crudely classify the word spans as construction
domain or out-of-domain, taking into account the span’s source corpus and frequency-based
characteristics. Sectio3.n4 explains how the domain terms are used to generate a KG, with
nodes automatically linked to Unicla2s8s] [and Wikidata4[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] – a large general domain
knowledge base. Sectio3n.5 describes various node-similarity measures and properties that
we compute between spans in the KG. After adding these additional relations to the KG, we
identify closely related terms through querying the graph, as well as by relying on network
metrics like clustering coeficient, centrality, and degree. The resulting set of relations, and
overall relatedness of terms, makes it possible to present annotators with an overview of terms
that are likely to belong to the same topic.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Term extraction and domain classification</title>
        <p>In previous work, we developed a Shallow Parser for Regulatory texts (SPaR.t4x6t].)S[PaR.txt
is originally trained to discover and identify terms – either single words or MWEs expressing a
single unit of information – in the Scottish Building regula4t7io].nWs[e run SPaR.txt over our
corpora and limit the identification of Object spans that express noun-based phrases, e.g.,
realworld objects or otherwise distinguishable concepts. Examples include ‘the fire-fighting’ alinftd
‘storage building’, but alsod‘esign process’ and ‘theory’. We post-process the SPaR.txt results
relatively strictly, to improve the quality of terms presented to human annotato1rs. Table
provides an overview of the numbers of candidate terms extracted from both sets of texts.
Inspired by [16] we distinguish domain from out-of-domain concepts based on the frequency
that a term was found in the foreground corpus and background corpus. We adopt a modified
Term Frequency-Inverse Document Frequency (TF-IDF) metric:
 -  ()
= (1 +</p>
        <p>+  
) ∗ (  
)

(1)
with   the number of times ter moccurs in the foreground corpu s,  the background corpus
count, an d</p>
        <p>the averaged IDF weight over the subword tokens of te.rm</p>
        <p>
          We compute term embeddings using a basic case-sensitive pre-trained BERT language model
[
          <xref ref-type="bibr" rid="ref34">48</xref>
          ], which we multiply with the average IDF-weight of a term’s tokens and normalise as
suggested by 3[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. We compute the Nearest Neighbours (NN) graph to identify the 500 most
similar embeddings for each term, relying on cosine similarity. Fi2gpulreots the number of
NNs that only appear in the foreground corpus against the modified TF-IDF value for terms.
Note that our TF-IDF value defaults to 0 if a term only occurs in the background corpus. Terms
are labelled as out-of-domain when both (1) their TF-IDF value is b0e.6lo,awnd (2) less than
200 out of the500 NNs can be found only in the foreground corpus. This labelling strategy splits
the total 7,940 terms into 4,958 domain and 2,982 general/out-of-domain terms. Note that these
numbers depend highly on the size of the input data, cleaning, and the domain-classification
values that we picked.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Building the KG</title>
        <p>We build our KG as a directed edge-labelled RDF graph. We distinguish between (1) the 295
defined terms that were manually extracted from MAD and (2) spans of text. The former
are added as instances of ‘skos:Conce’,patnd provided with their preferred label, alternative
labels and definitions as found in MAD. The latter are currently addierdeca:Csh‘aracterSpan’2
instances; these include terms that were automatically extracted from text, terms found in the
index term lists of MAD and terms found in our manually assembled KG. If a span was identified
by SPaR.txt, this is indicated using ‘prov:wasAttributed’T.o</p>
        <p>Defined concepts and spans are kept in separate namespaces, the corresponding nodes are
linked through ‘skos:exactMatc’h. We add a Uniclass namespace to store Uniclass terms that
occur verbatim in MAD, again linking them to a span node that has the same ‘rd’f–s:lwabeel
create this span if it did not exist in the KG yet. We also add a Wikidata namespace, to which
we add nodes for those concepts and spans that can be found in Wikidata. The source of nodes
is annotated using ‘prov:hasPrimarySour’c.e</p>
        <p>Where available, we add Wikidata definitions and classification labels to the span nodes in
our graph. Wikidata contains much noise, and many out-of-domain interpretations – e.g., the
span ‘wall’ is amongst others identified as a ‘Unix utili’t.yWe manually annotate the relevance
of 1,220 distinct Wikidata classes that are linked to the defined and index terms from MAD,
and find that 45.7% of these are irrelevant. We only keep definitions and corresponding class
labels, if the label is within our set of classes that were annotated as relevant. We currently add
definitions using ‘irec:wikiDefinition ’ and add class labels as separate spans that are linked with
‘rdf:type’. We also provide an indication of whether a term may be specific to the CC domain or
not, based on the domain classification described in Sec3t.i3o.n</p>
        <p>
          For the defined terms found in MAD, 25% of the 295 preferred and alternative labels occur
in Wikidata. And while 29% of our spans can be found in Wikidata, this number drops to
13% after filtering out irrelevant Wikidata classes. This indicates that Wikidata (besides being
very noisy), as expected, does not provide the coverage of building domain terms that would be
needed for CC. Finally, we run SPaR.txt on each of the Wikidata definitions. The aim is to
identify whether defined spans are related, based on overlapping terms in their definitions –
inspired by [
          <xref ref-type="bibr" rid="ref35">49</xref>
          ]. From the definitions we identify an additional 2,741 spans, which are added to
the graph and related to the respective defined spans with ‘irec:definitionRela’t.ion
2We define the CharacterSpan class and various properties in our own ’‘inreacmespace, some of these may be
substituted with more commonly used resources to improve interoperability.
        </p>
        <p>MAD Uniclass Wikidata
0 0 90
0 18 93
90 93 425
Nr. of defined concepts by source
MAD Uniclass WikiData
295 27 1,218
0 571 32</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.5. Identifying domain concepts</title>
        <p>The KG contains 2.1K concepts with 1.6K unique preferred labels. T2apbrloevides some insight
in the number of concepts, their labels and their definitions. Seven labels have both a concept
node with definition and a concept node without a definition; an example is ‘floating flo’.or
Notably, we do not include any definitions from Uniclass into our KG.</p>
        <p>
          There exist several indicators for spans in the KG to be a good candidate for a CC
conceptualisation. In the first place, one could rely on the presence of an exact match to a concept
from MAD, Uniclass or Wikidata. Besides a span being represented by one or more concepts,
one could consider looking at closely related spans. To this end, we compute several span-span
features, such as whether a span occurs in the definition of another span – inspired49b].y [
This means we run SPaR.txt over all definitions and identify a set of 2.7K additional spans.
Other computed features include:
• Morphological similarity, e.g., based on word overlap and edit distance we determine that
‘structural element’ is morphologically similar to ‘element of struct’u. re
• Semantic similarity, we embed spans and the 5 NNs for each span.
• Domain classification, we rely on our crude domain specificity – see Secti3o.3n–
classification to assign a label to spans that were not seen in our corporwa,aete.grp.,r‘oofing
membrane’ occurs in the definition for ‘green roof’ and is classified as CC domain.
• We identify potential acronyms in MAD, e.LgP.,A‘’ stands for ‘local planning authority’.
• Antonym-based features, based on the 3.3K antonyms found in WordNet [
          <xref ref-type="bibr" rid="ref20">34</xref>
          ].
Nodes in our KG that are highly connected, are likely to include inflections, alternative labels
and sub/super classes. We explore using the Louvain method for community detec5t0]iotno[
identify highly connected groups of nodes. To this end we convert the KG to a weighted graph
with span-labels as nodes, where edge weights are manually set based on the relation types.
This makes it possible to compute and visualise spans that are highly related in the KG, see
Figure 3 for an example.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Initial annotator feedback</title>
      <p>Together with the annotators we used a visual interface of the graph, as well as the graph
metrics computed before, to:
• Explore a new subtopic to define for an CC conceptualisation.
• Compare the terms in an existing subtopic against those captured in a graph, to see if
there are additional terms that we might want to add.</p>
      <p>
        First, the new subtopic we aim to map is terminology revolving around ‘venti’l.aFtiognure 4
visualises a part of the terms in the KG revolving around ventilation. From the KG and graphs
like the one shown in Figur3e, annotators quickly identify a hierarchy of high-level terms,
along with some subclasses and definitions. As an example, the relatively high level te’rm ‘flue
has subclassesv‘ent’, ‘duct’, and ‘chimney’. However, our annotators were asked to rely on the
same spreadsheet approach as used in the manual exploration of building a KG– see ApAp.endix
They find that working in a spreadsheet severely limits the types of relations they would like to
add to the KG. As an example, while a ‘cable d’umctay be a type of ‘duc’t, but it falls outside of
the scope of ventilation-related terms. Similarly, ventilation terminology is interrelated with
drainage terminology. At this stage, it may be beneficial to work with vocabulary editors, such
as VocBench 3 [
        <xref ref-type="bibr" rid="ref37">51</xref>
        ].
      </p>
      <p>We compare the manually assembled terms revolving arotuhnedrm‘al insulation’ against
those found in our KG. Our manual approach required two days of manual annotation which
resulted in a hierarchy of 30 concepts, with a total of 40 alternative labels, 19 definitions and 13
links to concepts in external resources. Using our KG approach, the annotators were able to
identify 15 additional terms that hadn’t been considered before in ten minutes. These included
missing classes at various depths in the hierarchy, as well as alternative labels. For the new and
existing concepts additional definitions from Wikidata were found, although these are often
too generic – e.g., ‘thermal insulation’(Q918306) is defined as ‘insulation against heat transfer ’.</p>
      <p>Overall, our annotators appreciate the graphical overviews and having terms, definitions and
source information gathered together. They find that the related terms are grouped together
quite well, which makes it easier to get a comprehensive overview of alternative and related
terms for a concept. The access to various surface forms (inflections) is more useful from a
computing perspective. However, while adding new terminology following our spreadsheet
approach – see AppendixA – the annotators stress the need for a better editing environment.
The consensus is that, ideally, the graphical overview of the KG allows adding new concepts
and relations directly.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and limitations</title>
      <p>Solving ACC touches on a variety of open research problems, such as semantic parsing,
canonicalisation and ontology matching. We present work on one of these sub-tasks, namely identifying
the lexicon of terms that may be used to compose CC rules. Candidate terms are for a large
part directly extracted from building regulations. By adding the terms to a KG, it is possible
to capture a large variety of relations between the candidates, as well as link them to external
resources. External resources that we relied on are an example of an existing vocabulary within
the building domain, Uniclass, and an example of a large general domain knowledge base,
Wikidata. We show that (1) regulatory texts can provide a basis for a CC conceptualisation,
and (2) using a KG to capture terminology can serve as a promising support tool for developing
such a conceptualisation.</p>
      <p>Our approach is exploratory and should be seen as a proof-of-concept. Most of the computing
steps may be tweaked, or replaced. As an example, the way we currently compute some features,
such as morphological similarity between sp(a ns2), limits the scalability of our approach.
Further research would be required to determine, e.g., the value of computed features, the
relevance of Wikidata classes, the domain classification. As noted, our KG schema contains
several idiosyncratic RDF classes and properties that may be replaced with more widely used
equivalents. Furthermore, due to licensing restrictions our KG only covers the MAD. This limits
the scope and usability of the KG, as well as its use for suggesting terms that should be included
in a CC lexicon.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This research is part of the intelligent Regulatory Compliance (i-ReC) project, a collaboration
between Northumbria University and Heriot-Watt University. We are grateful to the Building
Research Establishment (BRE), the Construction Innovation Hub (CIH), as well as Northumbria
University for funding this research. Our thanks also go to Ian Babelon and Huyam Abudib for
their help with annotation.
[13] F. Ameri, B. Kulvatunyou, N. Ivezic, K. Kaikhah, Ontological conceptualization based
on the SKOS, Journal of Computing and Information Science in Engineering 14 (2014).
doi:10.1115/1.4027582.
[14] S. Sarawagi, Information Extraction, Foundations and Trends® in Databases 1 (2007)
261–377. URL: http://pages.cs.wisc.edu/~anhai/courses/784-fall13/ieSurvey.pdf.1d0o.i:
1561/1500000003.
[15] M. Palmer, D. Gildea, N. Xue, Semantic role labeling, volume 3, 20101.0d.o2i2: 00/</p>
      <p>S00239ED1V01Y200912HLT006.
[16] A. Meyers, Y. He, Z. Glass, J. Ortega, S. Liao, A. Grieve-Smith, R. Grishman, O.
BabkoMalaya, The termolator: Terminology recognition based on chunking, statistical and
search-based scores, Frontiers in Research Metrics and Analytics 3:19 (2018) 1–1140..doi:
3389/FRMA.2018.00019/FULL.
[17] E. Hjelseth, N. Nisbet, Capturing Normative Constraints By Use of the Semantic Mark-Up
Rase, in: Proceedings of CIB, March, 2011, pp. 26–28. URLh:ttp://itc.scix.net/data/works/
att/w78-2011-Paper-45.p d.f
[18] J. Zhang, N. M. El-Gohary, Semantic NLP-Based Information Extraction from Construction
Regulatory Documents for Automated Compliance Checking, Journal of Computing in
Civil Engineering 30 (2016) 04015014. do1i:0.1061/(asce)cp.1943-5487.0000346.
[19] R. Zhang, N. El-Gohary, A Machine-Learning Approach for Semantically-Enriched
Building-Code Sentence Generation for Automatic Semantic Analysis, in: Construction
Research Congress, 2020, pp. 1261–1270.
[20] J. M. Siskind, A computational study of cross-situational techniques for learning
word-tomeaning mappings, Cognition 61 (1996) 39–91. do1i0:.1016/s0010-0277(96)00728-7.
[21] I. A. Sag, T. Baldwin, F. Bond, A. Copestake, D. Flickinger, Multiword expressions: A pain
in the neck for NLP, in: Lecture Notes in Computer Science (including subseries Lecture
Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), volume 2276, 2002,
pp. 1–15. doi:10.1007/3-540-45715-1_1.
[22] C. Ramisch, S. R. Cordeiro, A. Savary, V. Vincze, V. B. Mititelu, A. Bhatia, M. Buljan, M.
Candito, P. Gantar, V. Giouli, T. Güngör, A. Hawwari, U. Iñurrieta, J. Kovalevskaite, S. Krek,
T. Lichte, C. Liebeskind, J. Monti, C. P. Escartín, B. QasemiZadeh, R. Ramisch, N. Schneider,
I. Stoyanova, A. Vaidya, A. Walsh, Edition 1.1 of the Parseme shared task on automatic
identification of verbal multiword expressions, in: LAW-MWE-CxG 2018 - Joint Workshop on
Linguistic Annotation, Multiword Expressions and Constructions, Proceedings of the
Workshop, 2018, pp. 222–240. URL: https://gitlab.com/parseme/sharedtask-guidelines/issues.
[23] T. Baldwin, S. N. Kim, Multiword Expressions, in: N. Indurkhya, F. J. Damerau (Eds.),
Handbook of Natural Language Processing, second ed., Chapman and Hall, 2010, pp.
267–292.
[24] L. Getoor, A. Machanavajjhala, Entity resolution, Proceedings of the VLDB Endowment 5
(2012) 2018–2019. URL: https://dl.acm.org/doi/10.14778/2367502.2367564. do1i0:.14778/
2367502.2367564.
[25] V. Christophides, V. Efthymiou, T. Palpanas, G. Papadakis, K. Stefanidis, An Overview of
End-to-End Entity Resolution for Big Data, ACM Computing Surveys 53 (20211).0d.oi:
1145/3418896.
[26] L. Galárraga, G. Heitz, K. Murphy, F. M. Suchanek, Canonicalizing Open Knowledge Bases,</p>
    </sec>
    <sec id="sec-7">
      <title>A. Manually developing a KG for ACC</title>
      <p>
        We start by exploring the manual extension and linking of controlled vocabularies in the building
domain. While this approach is not scalable, the development can provide important lessons
for approaching the creation of a domain vocabula5r2y].[We model our vocabulary using the
Simple Knowledge Organisation System (SKOS), a common data model for sharing and linking
knowledge resources like thesauri, taxonomies and classification schemes [
        <xref ref-type="bibr" rid="ref39">53</xref>
        ].
      </p>
      <p>An example of a controlled vocabulary in the building domain is Unic2l8a].ssH[owever,
only 598 (4%) of the 15K Uniclass terms occur verbatim in the 1.274 pages of the UK Merged
Approved documents (MAD). This means that there is a severe mismatch between the wording
used in the regulations and the wording used in Uniclass. On the one hand, Uniclass has a
far wider coverage. On the other hand, considering the wide coverage of Uniclass one would
expect that many of its terms can be found in a relatively generic set of building regulations
like the Approved Documents. In conclusion, Uniclass can be expected to require substantial
extension or reformatting if it is to provide a shared conceptualisation that aligns with the
regulation texts.</p>
      <sec id="sec-7-1">
        <title>A.1. Approach</title>
        <p>
          Our first step is to develop a workflow for the collection and annotation of terminology that
works for our annotators. We then record new entities and relations through existing vocabulary
editors, such as VocBench 35[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. However, after several weeks and a multitude of meetings, we
settled on using a spreadsheet to capture three small subdomains of interest. S3eefoTrable
an example of the template we used. The aim is to ensure consistency in capturing concepts,
alternative labels, definitions, exact matches and other SKOS-based triples.
        </p>
      </sec>
      <sec id="sec-7-2">
        <title>A.2. Findings</title>
        <p>Developing a workflow and adding terms to the KG manually takes a tremendous amount of
time. With the workflow in place it is still challenging to identify synonyms and hyponyms of
terms. Our three annotators indicated to have spent at least two full days each on composing
their part of the graph. As a result of these six days of work, only 302 vocabulary terms
were identified with a total of 130 alternative labels. In 214 cases, a term was linked to the
unique identifier of a matching term in an external resource; Uniclass (49%), NRM3 (4%), and
a selection of British Standards vocabularies (49%). Note that terms and definitions from the
British Standards vocabularies cannot be shared due to licensing restrictions. Furthermore, the
six days of work does not include the development of a workflow for annotation or developing
scripts to integrate their separate annotations.</p>
        <p>
          Initially, annotators identified relevant terms through search engines and keyword search in
classification systems and standards. However, they found it dificult to identify appropriate
sources of domain terminology. The quality of the sources is not always easy to judge, and a
lack of definitions complicates the identification of relations between terms. Another hurdle is
determining whether the terms identified are actually used in the regulations. Classification
systems, and Uniclass in particular, are found to contain classes that are inconsistent with every
day language use in industry. To exemplify this, many leaf nodes in Uniclass are amalgamations
of properties, such as ‘fibre cement profiled sheet self-supporting cladding systems ’. This finding
corroborates the statement that classification schemes represent the needs of the issuing agencies,
rather than the needs of users [
          <xref ref-type="bibr" rid="ref40">54</xref>
          ].
        </p>
        <p>In the end annotators agreed on a standard set of sources to use, accepting the risk of missing
specialist sources for particular subject areas – which may include the more obscure terms that
would be useful to capture in the KG. Sources such as the British Standards vocabularies are noted
to be particularly valuable. In many cases definitions makes it easier to identify synonymous or
closely related terms. As such, definitions reduce the need for domain knowledge and make the
annotation process less error prone. Nevertheless, establishing which relationship should exist
between terms in the KG can be dificult.</p>
        <p>It is hard to determine when a comprehensive overview of terminology has been achieved.
Annotators found themselves disagreeing on which terms to include and how they relate. They
thought ‘skos:related’ is too vague, and tended to interpret ‘skos:bro’adaesra
‘rdfs:subClassOf ’ relation. The latter would imply that both terms are classes in the KG schema, just as
‘skos:Concept’ is a class.
ay ilp nd la s h e</p>
        <p>t e
la stc eh ro
t p
u n
con leba rem oud i-st ito</p>
        <p>h r
C l t p in la</p>
        <p>2
e .
d .2 .2
o .1 .1
C 1 1
a h u
tfre te slu ic it</p>
        <p>h ‑s
isun rop ist eh io v . u l , i
t t ie n n a tc</p>
        <p>a
g h s io ay rm ud ed .n
la tc ikn ta ilcp ca irte lta sco teh ro am itoa
rm ud ta m ap ta eop slta ily m pn fo slu
eTh rop ro fro fo t p in oP foa ito is in</p>
        <p>h r</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>McKechnie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shaaban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lockiey</surname>
          </string-name>
          ,
          <source>Computer Assisted Processing of Large Unstructured Document Sets: A Case Study in the Construction Industry, in: Proceedings of the ACM Symposium on Document Engineering</source>
          ,
          <year>2001</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Cook</surname>
          </string-name>
          ,
          <article-title>How legal drafting may be central to fire safety debate -</article-title>
          <source>BBC News</source>
          ,
          <year>2017</year>
          . URL: https://www.bbc.co.uk/news/uk-41049510http://www.bbc.co.uk/news/uk-41049.
          <fpage>510</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Meijer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Visscher</surname>
          </string-name>
          , L. Sheridan,
          <article-title>Building regulations in Europe Part I: A comparison of the systems of building control in eight European countries</article-title>
          , Delft University Press Science, Delft,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Niemeijer</surname>
          </string-name>
          , B. De Vries,
          <string-name>
            <given-names>J.</given-names>
            <surname>Beetz</surname>
          </string-name>
          ,
          <article-title>Freedom through constraints: User-oriented architectural design</article-title>
          ,
          <source>Advanced Engineering Informatics</source>
          <volume>28</volume>
          (
          <year>2014</year>
          )
          <fpage>28</fpage>
          -
          <lpage>361</lpage>
          .
          <year>0d</year>
          .
          <year>o1i0</year>
          :16/j.aei.
          <year>2013</year>
          .
          <volume>11</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Preidel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Borrmann</surname>
          </string-name>
          ,
          <article-title>BIM-based code compliance checking</article-title>
          ,
          <source>in: Building Information Modeling: Technology Foundations and Industry Practice</source>
          , Springer International Publishing,
          <year>2018</year>
          , pp.
          <fpage>367</fpage>
          -
          <lpage>381</lpage>
          . URLh:ttps://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -92862-
          <issue>3</issue>
          _
          <fpage>2</fpage>
          .2 doi:10.1007/978-3-
          <fpage>319</fpage>
          -92862-3{\_}
          <fpage>22</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dimyadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Amor</surname>
          </string-name>
          ,
          <source>Automated Building Code Compliance Checking</source>
          . Where is it at?,
          <source>in: Proceedings of the CIB World Building Congress 2013</source>
          and
          <string-name>
            <given-names>Architectural</given-names>
            <surname>Management</surname>
          </string-name>
          &amp;
          <article-title>Integrated Design and Delivery Solutions (AMIDDS</article-title>
          ),
          <volume>380</volume>
          ,
          <year>2013</year>
          , pp.
          <fpage>172</fpage>
          -
          <lpage>185</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N. O.</given-names>
            <surname>Nawari</surname>
          </string-name>
          , Automating Code Compliance Checking,
          <source>MDPI - Buildings</source>
          <volume>9</volume>
          (
          <year>2019</year>
          )
          <article-title>86</article-title>
          . URL: https://www.mdpi.com/2075-5309/9/4/86.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kruiper</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Konstas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sadeghineko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Watson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <surname>Don't Shoehorn</surname>
          </string-name>
          ,
          <article-title>but Link Compliance Checking Data (</article-title>
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dimyadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Clifton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Spearpoint</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Amor</surname>
          </string-name>
          ,
          <article-title>Computerizing Regulatory Knowledge for Building Engineering Design</article-title>
          ,
          <source>Journal of Computing in Civil Engineering</source>
          <volume>30</volume>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1061/(asce)cp.
          <fpage>1943</fpage>
          -
          <volume>5487</volume>
          .
          <fpage>0000572</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dimyadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Amor</surname>
          </string-name>
          , Automating Conventional Compliance Audit Processes,
          <source>in: 14th IFIP International Conference on Product Lifecycle Management (PLM)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>324</fpage>
          -
          <lpage>334</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dimyadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Amor</surname>
          </string-name>
          ,
          <article-title>BIM-based compliance audit requirements for building consent processing</article-title>
          ., in: J.
          <string-name>
            <surname>Karlshøj</surname>
          </string-name>
          , R. Scherer (Eds.), eWork and eBusiness in Architecture, Engineering and Construction, September, CRC Press,
          <year>2018</year>
          , pp.
          <fpage>465</fpage>
          -
          <lpage>471</lpage>
          . URhLt:tps: //www.taylorfrancis.com/books/9780429013652. do1i0:.
          <volume>1201</volume>
          /9780429506215.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          ,
          <article-title>Ontology Learning from Text : An Overview, Learning (</article-title>
          <year>2004</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . in
          <source>: Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management - CIKM '14</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1679</fpage>
          -
          <lpage>1688</lpage>
          .
          <year>d1o0i</year>
          .:
          <volume>1145</volume>
          /2661829. 2662073.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Rasmussen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pauwels</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Hviid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlshøj</surname>
          </string-name>
          ,
          <article-title>Proposing a Central AEC Ontology That Allows for Domain Specific Extensions</article-title>
          ,
          <source>in: Lean and Computing in Construction Congress - Volume 1: Proceedings of the Joint Conference on Computing in Construction, July</source>
          , Heriot-Watt University, Edinburgh,
          <year>2017</year>
          , pp.
          <fpage>237</fpage>
          -
          <lpage>244</lpage>
          . URhtLt:p://itc.scix.net/cgi-bin/ works/Show?_
          <source>id=lc3-2017-153. doi1:0</source>
          .24928/JC3-2017/0153.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gelder</surname>
          </string-name>
          ,
          <article-title>The principles of a classification system for BIM: Uniclass 2015</article-title>
          ,
          <source>Proceedings of the 49th International Conference of the Architectural Science Association</source>
          <volume>1</volume>
          (
          <year>2015</year>
          )
          <fpage>287</fpage>
          -
          <lpage>297</lpage>
          . URL: https://anzasca.net/wp-content/uploads/2015/12/028_Gelder_ASA2015.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Che</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          , T. Liu,
          <article-title>Learning Semantic Hierarchies via Word Embeddings, in: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics</article-title>
          , Stroudsburg, PA, USA,
          <year>2014</year>
          , pp.
          <fpage>1199</fpage>
          -
          <lpage>1209</lpage>
          . URLh:ttp://ir.hit.edu.cn/~car/papers/ acl14embedding.pdfhttp://aclweb.org/anthology/P14-11131.
          <year>0d</year>
          .
          <year>o3i</year>
          :
          <volume>115</volume>
          /v1/
          <fpage>P14</fpage>
          -1113.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [30]
          <string-name>
            <surname>T. K. Landauer</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          <string-name>
            <surname>Dutnais</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Fbltz</surname>
            , G. Pumas,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Kintsch</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Menn</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Streeter</surname>
          </string-name>
          ,
          <article-title>A Solution to Plato's Problem: The Latent Semantic Analysis Theory of Acquisition, Induction, and Representation of Knowledge, Psychological Review 1 (</article-title>
          <year>1997</year>
          )
          <fpage>211</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Turney</surname>
          </string-name>
          , Distributional Semantics Beyond Words:
          <article-title>Supervised Learning of Analogy and Paraphrase (</article-title>
          <year>2013</year>
          ). URLh:ttps://arxiv.org/pdf/1310.5042.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>W.</given-names>
            <surname>Timkey</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. van Schijndel</surname>
          </string-name>
          ,
          <article-title>All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Stroudsburg, PA, USA,
          <year>2021</year>
          , pp.
          <fpage>4527</fpage>
          -
          <lpage>4546</lpage>
          . UhRtLt:ps: //aclanthology.org/
          <year>2021</year>
          .emnlp-main.
          <volume>3</volume>
          .7d2oi:
          <fpage>10</fpage>
          .18653/v1/
          <year>2021</year>
          .emnlp-main.
          <volume>372</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>V.</given-names>
            <surname>Shwartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Dagan</surname>
          </string-name>
          ,
          <article-title>Improving Hypernymy Detection with an Integrated Path-based and</article-title>
          <string-name>
            <surname>Distributional Method</surname>
          </string-name>
          (
          <year>2016</year>
          ).hUtRtLp:s://arxiv.org/pdf/1603.06076. pdfhttp://arxiv.org/abs/1603.06076.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [34]
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>a. Miller, WordNet: a lexical database for English</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>38</volume>
          (
          <year>1995</year>
          )
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          . doi:
          <volume>10</volume>
          .1145/219717.219748.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>El-Gohary, a Semantic Similarity-Based Method for Semi-Automated Ifc Extension</article-title>
          ,
          <source>Proceedings of 5th International /11th Construction Specialty Conference</source>
          (
          <year>2015</year>
          ). doi:https://open.library.ubc.ca/cIRcle/collections/52660/items/1.0076395.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>El-Gohary</surname>
          </string-name>
          ,
          <article-title>An Automated Relationship Classification to Support SemiAutomated IFC Extension</article-title>
          ,
          <source>in: Construction Research Congress</source>
          <year>2016</year>
          , American Society of Civil Engineers, Reston,
          <string-name>
            <surname>VA</surname>
          </string-name>
          ,
          <year>2016</year>
          , pp.
          <fpage>2039</fpage>
          -
          <lpage>2049</lpage>
          . URLh:ttp://ascelibrary.org/doi/10.1061/ 9780784479827.203. doi:
          <volume>10</volume>
          .1061/9780784479827.203.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>J.</given-names>
            <surname>Raad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wouter</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. van Harmelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Pernelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fatiha</surname>
          </string-name>
          ,
          <article-title>Detecting erroneous identity links on the web using network metrics</article-title>
          ,
          <source>in: The Semantic Web - ISWC</source>
          <year>2018</year>
          , volume
          <volume>11136</volume>
          of Lecture Notes in Computer Science, Springer International Publishing, Cham,
          <year>2018</year>
          , pp.
          <fpage>217</fpage>
          -
          <lpage>232</lpage>
          . URL: http://dx.doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -00671-6_13http://link.springer.com/ 10.1007/978-3-
          <fpage>030</fpage>
          -00671-6. doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -00671-6.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>HM</given-names>
            <surname>Government</surname>
          </string-name>
          ,
          <source>The Building Regulations</source>
          <year>2010</year>
          :
          <article-title>The merged approved documents</article-title>
          ,
          <source>Technical Report June</source>
          ,
          <year>2022</year>
          . URLh:ttps://assets.publishing.service.gov.uk/government/ uploads/system/uploads/attachment_data/file/1082748/Merged_Approved_Documents_ _
          <fpage>Jun2022</fpage>
          _.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>C. o. t. E. U.</given-names>
            <surname>European Parliament</surname>
          </string-name>
          ,
          <source>Council Directive</source>
          <volume>90</volume>
          /385/EEC - EN,
          <year>1990</year>
          . UhRtLt:p: //eur-lex.europa.eu/LexUriServ/LexUriServ.do?uri=CELEX:31990L0385:en:HTM. L
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>C. o. t. E. U.</given-names>
            <surname>European Parliament</surname>
          </string-name>
          ,
          <source>Council Directive</source>
          <volume>93</volume>
          /42/EEC- EN,
          <year>1993</year>
          . UhRtLt:ps: //eur-lex.europa.eu/LexUriServ/LexUriServ.do?uri=CELEX:
          <article-title>31993L0042:EN:HTM</article-title>
          .L
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>C. o. t. E. U.</given-names>
            <surname>European Parliament</surname>
          </string-name>
          , EU Directive 98/79/EC - EN,
          <year>1998</year>
          . URhLt:tps://eur-lex. europa.eu/legal-content/EN/TXT/HTML/?uri=
          <source>CELEX:31998L0079&amp;from=EN.</source>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>C. o. t. E. U.</given-names>
            <surname>European Parliament</surname>
          </string-name>
          ,
          <source>Regulation (EU)</source>
          <year>2017</year>
          /
          <fpage>745</fpage>
          - EN,
          <year>2017</year>
          . UhRtLt:ps: //eur-lex.europa.eu/eli/reg/2017/745/oj.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>C. o. t. E. U.</given-names>
            <surname>European Parliament</surname>
          </string-name>
          ,
          <source>Regulation (EU)</source>
          <year>2017</year>
          /
          <fpage>746</fpage>
          - EN,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [44]
          <string-name>
            <surname>Gwang-Yoon</surname>
            <given-names>Goh</given-names>
          </string-name>
          ,
          <article-title>Choosing a Reference Corpus for Keyword Calculation</article-title>
          ,
          <source>Linguistic Research</source>
          <volume>28</volume>
          (
          <year>2011</year>
          )
          <fpage>239</fpage>
          -
          <lpage>256</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .17250/khisli.28.1.201104.013.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vrandečić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krötzsch</surname>
          </string-name>
          ,
          <article-title>Wikidata: A free collaborative knowledgebase</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>57</volume>
          (
          <year>2014</year>
          )
          <fpage>78</fpage>
          -
          <lpage>85</lpage>
          .
          <year>do1i</year>
          :
          <fpage>0</fpage>
          .1145/2629489.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>R.</given-names>
            <surname>Kruiper</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Konstas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sadeghineko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Watson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>SPaR.txt, a Cheap Shallow Parsing Approach for Regulatory Texts (</article-title>
          <year>2021</year>
          )
          <fpage>129</fpage>
          -
          <lpage>143</lpage>
          .
          <year>1d0o</year>
          .
          <year>i1</year>
          :
          <volume>8653</volume>
          /v1/
          <year>2021</year>
          . nllp-
          <volume>1</volume>
          .
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [47]
          <string-name>
            <surname>Scottish</surname>
            <given-names>Government</given-names>
          </string-name>
          ,
          <source>Building Standards Technical handbook 2020: Domestic</source>
          ,
          <year>2020</year>
          . URL: https://www.gov.scot/publications/ building-standards
          <article-title>-technical-handbook-2020-dome</article-title>
          .stic/
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          , in: arXiv preprint arXiv:
          <year>1810</year>
          .04805,
          <year>2018</year>
          . URL: https://github.com/tensorflow/tensor2tensorhttp://arxiv.org/abs/
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ko</surname>
          </string-name>
          ,
          <article-title>Query Reformulation for Descriptive Queries of Jargon Words Using a Knowledge Graph based on a Dictionary</article-title>
          ,
          <source>in: International Conference on Information and Knowledge Management, Proceedings, Association for Computing Machinery</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>854</fpage>
          -
          <lpage>862</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .1145/3459637.3482382.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>V. D.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Guillaume</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lambiotte</surname>
          </string-name>
          , E. Lefebvre,
          <article-title>Fast unfolding of communities in large networks</article-title>
          ,
          <source>Journal of Statistical Mechanics: Theory and Experiment</source>
          <year>2008</year>
          (
          <year>2008</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . doi:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -
          <lpage>5468</lpage>
          /
          <year>2008</year>
          /10/P10008.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stellato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fiorelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Turbati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lorenzetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. Van</given-names>
            <surname>Gemert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dechandon</surname>
          </string-name>
          , C. LaaboudiSpoiden,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gerencsér</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Waniart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Costetchi</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Keizer,</surname>
          </string-name>
          <article-title>VocBench 3: A collaborative Semantic Web editor for ontologies, thesauri and lexicons</article-title>
          ,
          <source>Semantic Web</source>
          <volume>11</volume>
          (
          <year>2020</year>
          )
          <fpage>855</fpage>
          -
          <lpage>881</lpage>
          . doi:
          <volume>10</volume>
          .3233/SW-200370.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Guha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brickley</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. MacBeth</surname>
          </string-name>
          , Schema.
          <source>org: Evolution of Structured Data on the Web, Queue</source>
          <volume>13</volume>
          (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1145/2857274.2857276.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>A.</given-names>
            <surname>Miles</surname>
          </string-name>
          , S. Bechhofer,
          <source>SKOS Simple Knowledge Organization System Reference</source>
          ,
          <year>2009</year>
          . URL: https://www.w3.org/TR/skos-reference/http://www.w3.org/TR/skos-reference/.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Cheng</surname>
          </string-name>
          , G. T. Lau,
          <string-name>
            <given-names>K. H.</given-names>
            <surname>Law</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <article-title>Regulation retrieval using industry specific taxonomies</article-title>
          ,
          <source>Artificial Intelligence and Law</source>
          <volume>16</volume>
          (
          <year>2008</year>
          )
          <fpage>277</fpage>
          -
          <lpage>303</lpage>
          .
          <year>d1o0i</year>
          .:
          <volume>1007</volume>
          / s10506-008-9065-5.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>