<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Accounting AI Measures as ISO/IEC 25000 Standards Measures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Trenta</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>UNINFO UNI TC</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Technical Committee Artificial Intelligence</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Turin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
      </contrib-group>
      <fpage>22</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>This paper is a part of a set of papers showing how newly defined data and software quality measures can be described in ISO 25000 format. In the first group of papers [3], [1], [28], [2], we discussed with the help of some examples, the general approach of conformance when new quality measures are defined, and in the last paper [20] how to build practical ISO/IEC 25000 compliant product quality measures for AI, starting from measures developed in several public projects. In this paper we continue to show, through some examples, that standards and research coming from the scientific community on the topic of AI measures can be easily accounted as ISO/IEC 25000 measures. Moreover, the paper can be considered for the works in AI standardization area.</p>
      </abstract>
      <kwd-group>
        <kwd>1 product quality</kwd>
        <kwd>measures</kwd>
        <kwd>accuracy</kwd>
        <kwd>ISO</kwd>
        <kwd>ISO/IEC 25059</kwd>
        <kwd>ISO/IEC 5259</kwd>
        <kwd>ISO/IEC 24029</kwd>
        <kwd>ITU-T F</kwd>
        <kwd>748</kwd>
        <kwd>11</kwd>
        <kwd>metric</kwd>
        <kwd>AI</kwd>
        <kwd>ML</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Artificial Intelligence</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Policy makers, industries, and academia are
facing the problem of building trust in AI; in this
paper we show, following the approach of a
previous paper [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], how some AI measures taken
from non-ISO standards and research literature,
can be accounted as ISO/IEC 25000 AI product
quality measures.
      </p>
      <p>The items considered for AI product quality
measures are recalled in the following “shopping
list”.</p>
      <p>
        For the following, it is useful to recall
definitions given in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>The implementation I is defined as a function
I= I(method, algorithm(library,
parameters), training(dataset, process))
of
where:
2 Note: for ‘algorithm’ it is intended the categorization of the code
that perform the task, e.g. for the classification task, the ‘algorithm’
can be either a neural network, or a decision tree, or a support
vector machine, or other.
2)</p>
      <p>Mij=Mij(I)
and taking into account 1):
3)</p>
      <p>Mij=Mij (method, algorithm(library,
parameters), training(dataset, process))
With those definitions, benchmark Bij is the
best value Mij for the time being (e.g. for a full
year) for the i-characteristic and the j-measure3
among all the K implementations of Ik</p>
      <p>Starting from those definitions, we map some
existing measures to ISO 25000 measures when
1) holds. In the following, we pick those existing
measures from</p>
      <sec id="sec-1-1">
        <title>A. ROC curve metric [24]</title>
        <p>
          B. Recommendation ITU-T F.748.11 [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ],
C. Holistic Evaluation of Language Models [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]
and explain how they can be accounted as
ISO/IEC 25000 measures and make some more
consideration about the perspectives of the
ongoing standardization work in the relevant
bodies on the topic of AI product4 evaluation and
assessment.
2. AI Standardization
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>UPDATE) (2023</title>
      <p>
        Policy makers have addressed the issue of AI
trustworthiness mainly, but not only, to the
international standardization body ISO/IEC SC42
and to the European standardization body
CEN/CENELEC JTC21 that have in charge the
drafting of technical standards in support of
industry and of lawful rules. For the scope of this
paper, we consider, among the others, the
standards based on ISO 25000 series that define
or contribute to define product quality for an AI
product [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The assessment of product quality,
possibly together with the assessment of process
quality [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], will be performed in the near future on
voluntary or mandatory basis, in the former case
to promote trustworthiness in AI systems, in the
latter case to get compliance to rules [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In the
following, we focus on ML based AI systems [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>The work of ISO\IEC SC42 in the last years
has given birth to a set of standards on AI that are
covering topics such as quality, testing, risk,
management system, data, application according
to the non-official scheme of figure 2.</p>
      <p>
        It is to be noted that SC42 has developed and
is developing extensions [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] to standards of
the series ISO/IEC 25000 and this appears at the
moment the most mature approach to the AI
product evaluation, as it relies on the core
SQuaRE standards developed since 2008. Indeed,
the ISO/IEC 25000 itself foresees the possibility
to extend the model to specific technologies like
AI, through the definition of new characteristics
and new measures. This view and its reasons are
also well explained in the ISO/IEC news given in
https://www.iec.ch/blog/new-internationalstandard-ensuring-quality-ai-systems.
      </p>
      <p>
        At the moment, the ISO 25000 extensions for
AI are the technical specification for AI product
3 in 4) the j-measure is supposed as scalar; if the j-measure is a vector
or a matrix, the expression 4) should be adapted.
4 Note: the topic of product measurement is distinct from the topic of
the process measurement.
quality evaluation [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] that is under development,
and the quality model for AI already published [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
that is to be read in conjunction with [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>The considerations of this paper are supporting
the current set of ISO standards.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Example: AUC (Area</title>
    </sec>
    <sec id="sec-4">
      <title>Receiver Operating Curve)</title>
    </sec>
    <sec id="sec-5">
      <title>Under</title>
      <p>
        A receiver operating characteristic (ROC)
curve is a graphical method for displaying true
positive rates and false positive rates across
multiple thresholds from a binary classifier [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
5 For the scope of this paper, we don’t discuss the characteristic to
which the measure of table 1 is referred; as hypothesis, it could be
referred to Functional correctness.
curve (AUC) can be calculated. Higher AUCs
indicate more robust performance, ranging from 0
(worst) to 1 (best). Classifiers that perform no
better than chance will have an AUC of 0.5
AUC is an example of how statistical methods for
assessing NN can be accounted as quality ISO
25000 measure (Table 1).
      </p>
      <p>
        In conclusion, an AI measure like AUC well
known in scientific literature and classified
according to [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] into the category of statistical
methods, can be represented in an ISO/IEC 25000
format5.
4. Example: Rec. ITU-T F.748.11
The Rec. F.748.11 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], proposes, among the
others, metrics for AI applications.
      </p>
      <p>The approach is to define benchmarks, as the
history of processors evolution that includes both
architecture, clock, energy consumption, and
more parameters, has shown that it is impossible
to make comparisons without a common
challenging metric, like e.g. FLOP/s. At the same
manner, for each triplet composed by
• Application (e.g. Image classification,</p>
      <p>Speech recognition,..),
• Dataset (e.g. Imagenet, LibriSpeech,..),
• ML model (e.g. ResNet, DeepSpeech2,..),
it is defined the benchmark</p>
      <p>• Accuracy
with a specific metric for each triplet (e.g. Word
Error Rate for Speech recognition implemented
with DeepSpeech2 and tested against dataset
LibriSpeech).</p>
      <p>
        The ML model is further detailed with neuron
layers, input size and source code, e.g.
• ML model detailed (e.g. ResNet_50)
• ML model source code (e.g.
https://github.com/KaimingHe/deepresidual-networks)
This corresponds to the use case 1 Accuracy
showed in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and can be represented as in Table
2 for the triplet Image Classification, Imagenet,
ResNet, implemented with ResNet_50 with
source code
https://github.com/KaimingHe/deepresidual-networks
      </p>
      <p>
        Secondly, the characteristics of the model are
identified; new characteristics are introduced
(calibration, toxicity) that are not present in
models [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] but can be handled as ISO 25000
conforming [27].
      </p>
      <p>
        Thirdly, the measures contain the same
description used in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]:
      </p>
      <sec id="sec-5-1">
        <title>With those considerations, it is easy to</title>
        <p>identify the full description of the measures
according the ISO/IEC 25000 format.</p>
        <p>
          For example, we consider the measure of
detection of toxic text6 and draw the table 3
below.
So, we can conclude that ITU-T F.748.11 [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]
measures can be accounted as ISO/IEC 25000
conforming measures.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Example: LLM</title>
    </sec>
    <sec id="sec-7">
      <title>Models) (Large Language</title>
      <p>
        In this clause we try to show how the measures
performed in the research [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] can be accounted
as ISO/IEC 25000 measures. To do this, we note
that in [60], the following criteria are applied, that
are the same criteria used for defining ISO 25000
compliant measures [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        Firstly, [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] taxonomizes the LLM
applications, as proposed in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
In conclusion, also the measures 7 for LLM
presented in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] can be mapped to ISO 25000
quality measure.
      </p>
    </sec>
    <sec id="sec-8">
      <title>6. Measurements in Operation</title>
      <p>
        As highlighted in [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], when an ML
implemented with neural networks uses
continuous learning, its hyperparameters are
evolving, and the measurement of characteristics
of the NN can be different (and assessed worse or
better) from the measurement taken in the initial
state. This is also the reason why the AI medical
devices are deployed and sold as “frozen” [26],
giving a guarantee to the user-buyer that the
behaviour of the ML will be the same all the time.
      </p>
      <p>Anyway, additional requirements (e.g.
operational performance not worse than tested
ones) and measurement can be satisfied, so
6 In NLP applications, there is the general task of text classification,
and among them there is the specific task for the machine to detect
prompts with toxic text (e.g.. biased questions, hate speech,…)
7 For the scope of this paper, we don’t discuss the characteristic to
which the measure of table 3 is referred; as hypothesis, it could be
referred to Functional correctness.
enlarging the field of evaluation, both along the
time and the post-training data and perform a
further assessment of the ML in the operational
mode.</p>
    </sec>
    <sec id="sec-9">
      <title>7. Formal Methods</title>
      <p>
        It is to be noted the awareness of the scientific
community for the need of an a-priori guarantee
of the robustness of NN: many papers ([
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]) contain the word
“certification” or “formal guarantees” or
“verification” or “provably”, as they research the
proof of a target performance.
      </p>
      <p>To understand how the topic is presently
addressed and to complete the landscape of AI
measures, we recall the approach represented by
formal methods.</p>
      <p>
        According to the classification given in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ],
formal methods can successfully answer the
question whether or not, for a given Neural
Network, input and output (e.g. input image of
airplane and output label “airplane”), a
modification of the input leads to the same output
or a different one (e.g. input image of airplane
with noise, output label “helicopter”). This
question can be formulated as a formal
mathematical statement that is either true or false
for a given neural network and image.
      </p>
      <p>
        Based on the research that have proven that
for Neural Network using the linear piece-wise
activation function (ReLU), it is possible to
measure robustness in terms of lower bound of
minimal adversarial distortion for given input data
points [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The results for ReLU, were
successively extended to NNs with common
activation functions like sigmoid [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The results
of various research on this topic are summarized
in §6.2 [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] that supports formal methods as
engineering or quality evaluation of some NNs
and characteristics.
      </p>
      <p>
        Even if formal methods are at the moment
considered [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] a quality approach
complementary to ISO/IEC 25000, we could
consider the math function that defines distortion
bounds8 as a SQuaRE measurement function and
then account any formal method as an ISO/IEC
25000 measure.
      </p>
      <p>8 Formal methods are based on the theorem that, in certain
conditions, there exist Upper and Lower bounds for an m-layer
neural network function f with ! neurons so that for Ɐj ∈ [!] and
Ɐx ∈ ( | − "| ≤ ) holds:</p>
    </sec>
    <sec id="sec-10">
      <title>8. Proposal</title>
      <p>
        The proposal in this paper completes the
proposal in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]; there we showed how to design
and document a product quality measure that
includes algorithm, training dataset, library code
and parameters; here we show in a sort of reverse
engineering, how to account and represent
measures from standard and scientific literature
into the ISO/IEC 25000 format.
      </p>
      <p>
        Finally, some investigation areas (formal
methods and operational measures) and relevant
considerations are presented in the perspective of
an even wider application of the present and [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
paper proposals.
      </p>
    </sec>
    <sec id="sec-11">
      <title>9. Conclusion</title>
      <p>The role of ISO/IEC 25000 in measurement
and assessing of AI product quality is widely
recognized and of growing interest.</p>
      <p>
        A further confirmation comes from the
similarity between measurement methods
developed in scientific literature and projects and
the ISO 25000 conforming measurement method
as shown in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and in this paper.
      </p>
      <p>The paper analyzed this similarity and came to
the conclusion that most of the measurement
methods used for AI can be easily mapped into
ISO 25000 format.
10.
#$(x) ≤ #(x) ≤ #%(x)</p>
      <p>Where x is the perturbed input vector, centered in the reference
data point x0 and bounded in a sphere with ray ε.
the robustness of neural networks — Part 2:
Methodology for the use of formal methods
[26] M. van Hartskamp, S. Consoli er al.,
Artificial Intelligence in Clinical Health Care
Applications: Viewpoint, Interactive journal
of medical research, 2019
[27] International Organization for
Standardization, ISO/IEC DIS 25002
Systems and Software engineering - Systems
and software Quality Requirements and
Evaluation (SQuaRE) - Quality models
overview and usage
[28] A. Simonetta, A. Trenta, M. C. Paoletti, and
A. Vetrò, “Metrics for identifying bias in
datasets,” SYSTEM, 2021.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trenta</surname>
          </string-name>
          ,
          <article-title>Data bias measurement: a geometrical approach through frames</article-title>
          ,
          <source>Proceedings of IWESQ@APSEC</source>
          <year>2021</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3114</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trenta</surname>
          </string-name>
          <article-title>: ISO/IEC 25000 quality measures for A.I.: a geometrical approach</article-title>
          ,
          <source>Proceedings of IWESQ@APSEC</source>
          <year>2020</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2800</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Natale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Trenta</surname>
          </string-name>
          ,
          <source>Examples of practical use of ISO/IEC 25000, Proceedings of IWESQ@APSEC</source>
          <year>2019</year>
          . URL: http://ceurws.org/Vol-
          <volume>2545</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>International</given-names>
            <surname>Organization</surname>
          </string-name>
          for Standardization,
          <source>ISO/IEC 22989:2022 Information technology - Artificial intelligence -Artificial intelligence concepts</source>
          and
          <source>terminology</source>
          . URL: https://www.iso.org/standard/74296.html
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>International</given-names>
            <surname>Organization</surname>
          </string-name>
          for Standardization,
          <source>ISO/IEC DIS 25059 Software engineering - Systems and software Quality Requirements</source>
          and
          <string-name>
            <surname>Evaluation (SQuaRE</surname>
          </string-name>
          )
          <article-title>- Quality Model for AI-based systems</article-title>
          . URL: https://www.iso.org/standard/80655.html
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>International</given-names>
            <surname>Organization</surname>
          </string-name>
          for Standardization, ISO/IEC CD 5259
          <article-title>-2 (under development) Artificial intelligence - Data quality for analytics and ML - Part 2: Data quality measures</article-title>
          . URL: https://www.iso.org/standard/81860.html
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>International</given-names>
            <surname>Organization</surname>
          </string-name>
          for Standardization, ISO/IEC 23053:
          <year>2022</year>
          <article-title>Framework for Artificial Intelligence (AI) Systems Using Machine Learning (ML)</article-title>
          . URL: https://www.iso.org/standard/74438.html
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Natale</surname>
          </string-name>
          ,
          <article-title>Extensions of ISO/IEC 25000 quality models to the context of Artificial Intelligence</article-title>
          ,
          <source>Proceedings of IWESQ@APSEC 2022</source>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>International</given-names>
            <surname>Organization</surname>
          </string-name>
          for Standardization, ISO/IEC 42001 (
          <article-title>draft) Information technology - Artificial intelligence - Management system</article-title>
          . URL: https://www.iso.org/standard/81230.html
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>European</surname>
            <given-names>Commission</given-names>
          </string-name>
          , COM/
          <year>2021</year>
          /206 '
          <article-title>Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts</article-title>
          ',
          <year>2021</year>
          . URL: https://eur-lex.europa.eu/legalcontent/EN/TXT/?uri=
          <source>CELEX:52021PC020 6</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11] International Organization for Standardization,
          <source>ISO/IEC 23053:2022 Framework for Artificial Intelligence (AI</source>
          ) URL: https://www.iso.org/standard/74438.html
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hein</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Andriushchenko</surname>
          </string-name>
          , “
          <article-title>Formal guarantees on the robustness of a classifier against adversarial manipulation</article-title>
          ,” in NIPS,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Zhang</surname>
            <given-names>H.</given-names>
          </string-name>
          , Weng T.-W., Chen P.-Y.,
          <string-name>
            <surname>Hsieh C</surname>
          </string-name>
          .
          <article-title>-</article-title>
          J., Daniel L.
          <article-title>Efficient Neural Network Robustness Certification with General Activation Functions</article-title>
          .
          <source>Neural Information Processing Systems Conference</source>
          .
          <year>2018</year>
          ,
          <volume>31</volume>
          ,
          <fpage>4944</fpage>
          -
          <lpage>4953</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sinha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Namkoong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Duchi</surname>
          </string-name>
          , “
          <article-title>Certifiable distributional robustness with principled adversarial training</article-title>
          ,
          <source>” ICLR</source>
          ,
          <year>2018</year>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Raghunathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          , “
          <article-title>Certified defenses against adversarial examples</article-title>
          ,
          <source>”ICLR</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>T.-W. Weng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-J. Hsieh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Boning</surname>
            ,
            <given-names>I. S.</given-names>
          </string-name>
          <string-name>
            <surname>Dhillon</surname>
          </string-name>
          , and L. Daniel, “
          <article-title>Towards fast computation of certified robustness for relu networks</article-title>
          ,
          <source>” ICML</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T.</given-names>
            <surname>Gehr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mirman</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
            Drachsler-Cohen,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Tsankov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Chaudhuri</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Vechev</surname>
          </string-name>
          , “
          <article-title>Ai2: Safety and robustness certification of neural networks with abstract interpretation,”</article-title>
          <source>in IEEE Symposium on Security and Privacy (SP)</source>
          , vol.
          <volume>00</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>948</fpage>
          -
          <lpage>963</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Mirman</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gehr</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vechev</surname>
            <given-names>M</given-names>
          </string-name>
          .
          <article-title>Differentiable Abstract Interpretation for 1020 Provably Robust Neural Networks</article-title>
          .
          <source>Proceedings of the 35th International Conference on Machine Learning</source>
          .
          <year>2018</year>
          ,
          <volume>80</volume>
          ,
          <fpage>3575</fpage>
          -
          <lpage>3583</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bommasani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lee</surname>
          </string-name>
          et al.,
          <source>Holistic Evaluation of Language Models</source>
          ,
          <source>Stanford Institute for Human-Centered Artificial Intelligence (HAI)</source>
          , Stanford University,
          <year>2022</year>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Trenta</surname>
          </string-name>
          : ISO/IEC 25000 and
          <string-name>
            <given-names>AI</given-names>
            <surname>Product Quality Measurement Perspectives Proceedings APSEC IWESQ</surname>
          </string-name>
          <article-title>2022 (https://ceur-ws</article-title>
          .
          <source>org/</source>
          Vol-
          <volume>3356</volume>
          /, SSN 1613- 0073,https://dblp.org/db/conf/apsec/iwesq20 22.html#GirirajHH22)
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>ITU-T F</surname>
          </string-name>
          .
          <volume>748</volume>
          .
          <article-title>11 Metrics and evaluation methods for a deep neural network processor benchmark, 2020</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>ITU-T F</surname>
          </string-name>
          .
          <volume>748</volume>
          .
          <article-title>12 Deep learning software framework evaluation methodology, 2021</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <article-title>International Organization for Standardization, ISO/IEC DTS 25058 Software engineering - Systems and software Quality Requirements and Evaluation (SquaRE) - Guidance for quality evaluation of artificial intelligence (AI) systems</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24] International Organization for Standardization,
          <source>ISO/IEC TR 24029-1:2021 Artificial Intelligence (AI</source>
          )
          <article-title>- Assessment of the robustness of neural networks - Part 1: Overview</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25] International Organization for Standardization,
          <source>ISO/IEC 24029-2:2023 Artificial intelligence (AI</source>
          )
          <article-title>- Assessment of</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>