<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Novel Cephalometric Tool Enhanced by AI Assistance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Riccardo Zese</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Lombardo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo De Maio</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Tamascelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca Cremonini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DE, University of Ferrara</institution>
          ,
          <addr-line>Via Saragat 1 I-44122, Ferrara</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DMTR, University of Ferrara</institution>
          ,
          <addr-line>Via Luigi Borsari 46, I-44121, Ferrara</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>DOCPAS, University of Ferrara</institution>
          ,
          <addr-line>Via Luigi Borsari 46, I-44121, Ferrara</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Lateral radiography is one of the most important records for patients' evaluation in orthodontics and cephalometric analysis is fundamental to conduct correct diagnosis and treatment plan. This analysis includes both linear and angular measurements that quantitatively describe cranial and intermaxillary relationships. In order to obtain such measurements, anatomical landmarks are used. These reference points can be found on the soft tissue profile and on hard tissues such as teeth and skeletal contour. It is important to be extremely precise in the identification of these landmarks to compute correct measurements: even the slightest discrepancy could result in wrong values leading to diferent and possibly erroneous treatment plan. The automatic computerized identification of such anatomical landmarks on lateral cephalograms would greatly simplify this important step in the diagnostic process. Our aim is to apply artificial intelligence techniques for the automatic detection of these landmarks, with the final objective of developing a software, THERE (auTomatic HElpeR for cEphalometry), which exploits a predictive model that analyses teleradiographs, returns the coordinates of the anatomical landmarks, and automatically calculates the measurements necessary for diagnosis. This short paper describes the system interface and the first results obtained towards the training of the model(s) for landmarks prediction.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Machine learning</kwd>
        <kwd>Diagnostic methods and tools</kwd>
        <kwd>Odontology</kwd>
        <kwd>Artificial intelligence</kwd>
        <kwd>Computer vision</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Cephalometric analysis is of primary importance for the clinical evaluation of the orthodontic
patient: it is essential for a correct diagnosis and for choosing the right treatment plan. It is a
descriptive quantitative analysis, performed on a specific radiograph - called lateral
teleradiograph of the head - for all the patients that need to undergo an orthodontic treatment. This
patient’s profile radiograph must be carried out in maximum intercuspation and will result in a
1:1 image of the patient’s skull, analogically printed or digitally displayed. A set of anatomical
landmarks, which are reference points located both on hard and soft tissues of the profile, will
be identified on this image. The accuracy in the identification of these landmarks will be the
basis of the reliability of the measurements taken during the cephalometric analysis, i.e., the
more accurate the identification of these landmarks, the more correspondence there will be
with the real facial and dental pattern of the patient under analysis. The anatomical landmarks
on the radiographic image allow the computation of planes, defined by angles and lines, used to
describe and, if present, quantify dentoalveolar and skeletal anomalies.</p>
      <p>There are numerous anatomical landmarks that can be identified in a cephalometry, thus,
many diferent measures can be defined depending on which landmarks are considered. The
choice of these measures depends on the type of investigation to be conducted with the
cephalometry. For instance, the anatomical landmarks S (center of the sella turcica), N (most anterior point
of the fronto-nasal suture), and A (most posterior point of the maxillary anterior concavity)
define an angle that represents the sagittal position of the upper jaw in relation to the cranial
base, describing its normal or excessively anterior/posterior positioning. Similarly, the angle
defined by S, N and B (most posterior point of the mandibular anterior concavity), represents the
sagittal position of the mandible with respect to the cranial base. The diference ANB between
angles SNA and SNB represents the sagittal intermaxillary relationship. Normal values range in
the interval 2 ± 2 degrees, whereas it represents anomalies outside this range, skeletal Class II
if &gt; 4 degrees or a skeletal Class III if &lt; 0 degrees.</p>
      <p>This shows how the precise, accurate and consistent identification of the initial anatomical
landmarks, on which the analysis measurements are based, is of crucial importance to give the
orthodontist a correct and undistorted view of the patient’s cranial relationships. Historically,
cephalometric analysis was performed in pencil on transparent sheets placed on the radiograph,
process that requires a certain degree of efort on the part of the medics for calibrating the
image and tracing the necessary anatomical landmarks and structures. In recent years this
analysis is performed digitally with specific software that asks the physician only to position
the anatomical landmarks on the radiograph removing the burden of calculating angles and
lines. The remaining challenge in this area lies in the ability of cephalometric analysis software
to automatically detect and position the anatomical landmarks, avoiding the prior need to
manually set them, in order to make the work of the orthodontist more eficient.</p>
      <p>Thus, cephalometry is suitable to be performed with the use of automatic systems since it
is performed by analysing precise and well-located landmarks on the skull. The coordinates
returned by the system can be drawn on the teleradiography, allowing the clinician to check the
results and ensuring maximum transparency of the model results. This transparency ensures a
high explainability (or interpretability) of the model. However, in the literature there is a lack
of efective procedures for self-detection of cephalometric anatomical landmarks.</p>
      <p>
        Bulatova et al. in 2021 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] studied the accuracy and reliability of the ability of cephalometric
landmarks detection performed by a well-known commercial system from DDH Inc. This study
considered 16 landmarks taken from 110 teleradiographs of the skull. No significant diferences
were found for 12 out of 16 landmarks between the position automatically identified by the
system and those identified by trained human operators, empirically demonstrating that AI
is a promising tool to facilitate the performance of cephalometric analysis in routine clinical
practice and can speed up the analysis of large databases for research purposes. Similarly, Kim
et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] analysed a database of 2075 images to automatically identify the position of important
landmarks in cephalometric analysis, obtaining an anatomical diagnosis of subjects based on
the correct landmark placement in 88.43% of cases.
      </p>
      <p>In this short paper we present preliminary work to develop a system able to return the
coordinates of the landmarks and to calculate all the needed measurements for the correct classification
of the three classes of malocclusion: protrusion, retrusion or absence of malocclusion. This
allows the medics to have a powerful diagnostic tool on his side that speeds up the diagnosis.
The main objective of the project will therefore be the training of a model capable of correctly
identifying malocclusions and the development of a system, distributed as a Web application,
using this model, called THERE. The final system will allow the orthodontist to upload the
radiographs and obtain the cephalometry analysis, accompanied by the landmarks obtained,
the classification and the measurements. This will allow even less experienced orthodontists to
minimise errors, even in cases where landmarks identification appears more complex, e.g., due
to poor image quality. THERE will also allow the user to check the landmarks obtained and, if
necessary, correct them if wrongly placed. The implementation of THERE as a web application
allows for easy distribution of the system, which can be used by the user in any situation, easy
maintenance, as each update will be made directly available without any action on the part
of the user, ensuring maximum anonymisation of patients, as it will not require any type of
data other than the image of the cephalometry, and will allow expert users to correct wrong
predictions on the one hand, and the application to continuously collect new data in order to
constantly improve its accuracy on the other hand.</p>
    </sec>
    <sec id="sec-2">
      <title>2. THERE Web Application</title>
      <p>The application has been developed using the Flask Framework, written in Python. It is
designed to allow to focus on application-level business logic, without unnecessary ties to
specific deployment environments. The application is developed as a package and exploit Flas
Blueprints to be easily extensible in future. They allows the encapsulation of functionality, such
as views and templates, helping the adoption of a Model-View-Controller pattern.</p>
      <p>The workflow of the application is linear and has been kept simple with the aim of developing
an application which is straightforward to use for a user, who could be not proficient in
computer use. Moreover, since the application has to work with teleradiographs from patients,
the application must ensure the highest privacy. For this reason, no sensible data is collected
by the application. However, if the user decides to significantly change the coordinates of the
anatomical landmarks found, the application will save internally the image, which does not
contain sensible data of the patient, and the corrected coordinates in order to improve the
performance of the underlying neural network.</p>
      <p>Basically, the Web application is composed of three main views: Home, which is the landing
page of the URL of the application, from which the user must start using the application;
Calibration, the page where the user has to set the application to work on the uploaded
image; Dashboard, the operating page, where the uploaded image is analysed, the anatomical
landmarks are shown and the user can both modify the coordinates of the landmarks and read
the measurements necessary for the diagnosis.</p>
      <p>The complete workflow is depicted in Figure 1. The user visits the Home page and uploads
an image. The user will be directed to the Calibrate page, where the uploaded image is shown,</p>
      <p>Home: Upload an image
3 Dashboard: Now, the user has the opportunity to:</p>
      <p>Change coordinated of Repere points.</p>
      <p>The measurements are automatically
updated.</p>
      <p>TheThyiswaillllbowesutsoedsatvoeimcpharonvgeetshmeamdoed.el
Recalibrate the application
Upload a new image
1
2
3
and the user is called to indicate the scale of the image. Cephalometries include a ruler, used
to calculate the correct scale of the image. So, in the Calibrate page, the user has to draw a
ruler, as shown by 2 in Figure 1, to calibrate the application in the same way they should do
with standard cephalometries. Finally, the user will enter the main Dashboard, where they
can modify the anatomical landmarks and read all the important measurements they need
instantaneously.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Anatomical landmarks detection</title>
      <p>The current dataset is composed of 1732 images (PNG or JPG) of anonymized lateral
teleradiographs of the head of diferent patients. Images are taken using diferent devices and present
diferent levels of saturation, colour and brightness. Each image is labelled with the coordinates
in pixels of the 14 anatomical landmarks considered in this project, shown in Figure 2. Images
size spans between 2685× 2232 and 316× 224 pixels, most of them in RGB format. All
teleradiographs are taken with the patient looking to the right side of the picture, as shown in Figure 2.</p>
      <p>One of the biggest problems to be solved in analysing images is their dimensionality. The needs
of high accuracy in the coordinates detection of the anatomical landmarks is the main objective
we must pursue when modelling and training the underlying model of the application. Moreover,
the number of landmarks to be found and their proximity further increase the complexity of
the problem. Preliminary experiments have shown how the complexity of teleradiography and
S: Central point of the Sella Turcica.</p>
      <p>N: Deepest antero-posterior point of the nose-frontal suture.</p>
      <p>ANS: Most anterior bony point of the anterior nasal spine.</p>
      <p>PNS: Radiological point determined by the perpendicular drawn from the apex of the pterygopalatine
fossa to the bispinal plane.</p>
      <p>A: Most posterior point of the anterior concavity of the maxillary alveolar process.</p>
      <p>B: Most recessed point of the anterior concavity of the mandibular alveolar process.</p>
      <p>U1 root/U1 tip: Root apex/incisal edge of the upper central incisor.</p>
      <p>L1 root/L1 tip: Root apex/incisal edge of the lower central incisor.</p>
      <p>Pg: Most anterior point of the mental symphysis contour.</p>
      <p>GN: Lowest point of the mental symphysis contour.</p>
      <p>Go: Geometric point constructed at the point where the tangent to the ascending branch of the mandible
meets the plane of the mandible.</p>
      <p>
        Mesial: apex of the mesiovestibular cusp of the upper first molar.
the low variance of the coordinates values among the diferent examples tends to force the
model to compute an average position for the diferent landmarks, achieving low loss values
but always returning similar coordinates. In this regard, reducing the size of images too much
could be a tricky path to follow, because it tends to facilitate the occurrence of this problem. In
our first tests we followed two diferent approaches: (1) we divided the anatomical landmarks
in four subsets, containing landmarks used together when computing the skull measurements,
and trained four models specialized on a single subset, the final results will be obtained by the
ensemble of the four models; and (2) we applied bigger networks to the whole set of landmarks.
Setting 1. In this setting all the images where resized to 224x224 pixels, the size of the smallest
image, and transformed in RGB format. Data was augmented by flipping images and applying
rotation (± 10) and scale (50%-70%) both with a probability of 0.3. The pixels values were scaled
in the range [
        <xref ref-type="bibr" rid="ref1">-1,1</xref>
        ]. We built four networks, one for each subset of landmarks, namely (S,N,A,B),
(ANS, PNS, GN, Go), (U1 root, U1 tip, L1 root, L1 tip), and (Pg, Mesial). The backbone of
the network is a EficientNetB7 pretrained on Imagenet, which is followed by two Separable
Convolution layers with kernel size of 5x5 the first and 3x3 the latter, stride 1, and 8 filters. We
used Nadam as optimizer with a learning rate of 0.001, which is reduced by a factor 0.2 when a
tableau is reached, and early stopping. As loss we considered the mean squared error. The best
model for each group achieved a test loss values between 0.00045 and 0.00068. All the models
need 9MB to be saved in memory. Figure 3 shows some examples of the results obtained by
the ensemble of the four networks. On average, the distance between the actual and predicted
points is 6.3± 4.6 mm.
      </p>
      <p>
        Setting 2. Before training the model, data in the training set were pre-processed by converting
all the images to greyscale, normalizing the pixels values, and resizing them into 1000x1000
pixels. Then, data augmentation was performed by adding to the training set three new images
for each original image created by: flipping the original image, computing linear contrast
and adding gaussian blur with a probability of 80%, and rotating the original and the flipped
images of ± 15 degrees. To try to facilitate training, the coordinates have been standardised
in the range [
        <xref ref-type="bibr" rid="ref1">-1,1</xref>
        ]. We tried two diferent types of networks: (N1) a classical CNN composed
of five blocks applying convolution (with 32, 64, 128, 256, and 512 filters with kernel sizes
of 3x3 except for the first block with a kernel size of 5x5) and max pooling, followed by a
global average pooling and a dense layer of 256 neurons with ReLU activation function and
the output layer returning the coordinates of the 14 landmarks; (N2) a network composed
of 5 inception modules, shown in Figure 4, with 64, 64, 96, 96, 128 filers for the 1x1
convolution, 64, 64, 128, 256 filters for the 3x3 convolution and 32, 32, 64, 64, 128 filters for the
5x5 convolution. The network ends with a dense layer of 1024 neurons with ReLU activation
function and the output layer returning the coordinates of the 14 landmarks. We used Adam
as optimizer with a learning rate of 0.001 and early stopping. Moreover, the learning rate
is reduced in case of plateau by a factor 0.2. As
loss we considered the mean squared error. Both Input/Max pool.
models achieved a mean squared error computed
on the test set near 0.03. On average, the distance Conv 1x1 Conv 3x3 Conv 5x5 Max pool. 2x2
between the actual and predicted points computed Batch norm. Batch norm. Batch norm.
by N1 network is 24.8± 7.9 mm, while that of N2
network is 11.5± 5.4 mm. We observed that, in Concat.
general, the model based on inception (N2) was
less prone to return the same coordinates for each Figure 4: Inception block
image, as can be seen in Figures 5 (green cross for
label landmarks, red dots for predicted), but it is significantly bigger than the model without
inception blocks (more than 9 · 106 parameters - 103MB vs 5.6 · 106 parameters - 20MB).
In this paper we presented a Web application called THERE, that allows users to upload a
teleradiography which is analysed by an underlying neural network to locate the coordinates
of the 14 anatomical landmarks necessary to perform a cephalometry. Preliminary work on
the neural network shows that the localization of these landmarks is not trivial to perform
automatically due to the type of images used. The considered models tend to learn the average
position of each landmark instead of concentrating on the image itself, leading to the need of
an accurate investigation of how the images should be fed to the model and how the model
should be designed. Better results have been achieved by considering subsets of landmarks
separately. We are currently working on diferent models, from sequential CNNs to the more
complex transformer [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and YOLO-Pose [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] architectures. We also plan to study possible
masks to apply to the image or the possibility to apply patching, i.e., create sub figures from
the whole teleradiography and train the network on the sub figures instead of the entire
image. We also plan to perform a validation of the efectiveness and usability of the system
by administering a questionnaire to the users of the application, based on Post-Study System
Usability Questionnaire (PSSUQ) Version 3 template [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Finally, we plan to add functionalities
to the system, such as performing classification about other pathologies using intra-oral pictures
of patients. Acknowledgments This work is financed by "Bando Giovani anno 2022 per progetti
di ricerca finanziati con il contributo 5x1000 anno 2020".
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Bulatova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kusnoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Grace</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. P.</given-names>
            <surname>Tsay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Avenetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. J. C.</given-names>
            <surname>Sanchez</surname>
          </string-name>
          ,
          <article-title>Assessment of automatic cephalometric landmark identification using artificial intelligence</article-title>
          ,
          <source>Orthodontics &amp; Craniofacial Research</source>
          <volume>24</volume>
          (
          <year>2021</year>
          )
          <fpage>37</fpage>
          -
          <lpage>42</lpage>
          . doi:
          <volume>10</volume>
          .1111/ocr.12542.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Shim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Web-based fully automated cephalometric analysis by deep learning</article-title>
          ,
          <source>Computer Methods and Programs in Biomedicine</source>
          <volume>194</volume>
          (
          <year>2020</year>
          )
          <article-title>105513</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.cmpb.
          <year>2020</year>
          .
          <volume>105513</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          , in: I. Guyon, U. von Luxburg, S. Bengio,
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. V. N.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9</source>
          ,
          <year>2017</year>
          , Long Beach, CA, USA,
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Maji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nagori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mathew</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Poddar</surname>
          </string-name>
          ,
          <article-title>Yolo-pose: Enhancing YOLO for multi person pose estimation using object keypoint similarity loss</article-title>
          ,
          <source>in: IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops</source>
          <year>2022</year>
          , New Orleans, LA, USA, June 19-20,
          <year>2022</year>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>2636</fpage>
          -
          <lpage>2645</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPRW56347.
          <year>2022</year>
          .
          <volume>00297</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <article-title>IBM computer usability satisfaction questionnaires: Psychometric evaluation and instructions for use</article-title>
          ,
          <source>International Journal of Human-Computer Interaction</source>
          <volume>7</volume>
          (
          <year>1995</year>
          )
          <fpage>57</fpage>
          -
          <lpage>78</lpage>
          . doi:
          <volume>10</volume>
          .1080/10447319509526110.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sauro</surname>
          </string-name>
          ,
          <article-title>Revisiting the factor structure of the system usability scale</article-title>
          ,
          <source>Journal of Usability Studies</source>
          <volume>12</volume>
          (
          <year>2017</year>
          )
          <fpage>183</fpage>
          -
          <lpage>192</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>