<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modeling of Expert Estimation*</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Tver State Agriculture Academy</institution>
          ,
          <addr-line>7 Marshala Vasilevskogo Street (Sakharovo), Tver, 170026</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tver State Technical University</institution>
          ,
          <addr-line>22 Af. Nikitina Embankment, Tver, 170026</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The problem of collective decision-making by a group of experts is a crucial one in the theory and practice of decision-making. To obtain a high-quality collective estimation of an object or process, individual expert opinions should be coordinated. In this article, we use the absolute error of estimation (deviation of individual expert estimates from their arithmetic mean) and the relative error (absolute error divided by the arithmetic mean of estimates) as indicators of individual estimates consistency within a collective decision. We consider errors as random variables that fall under the normal probability distribution law. For the selected indicators, their variances, probability distribution densities, and confidence intervals (for averages and variances) are obtained. The preference criterion of methods for constructing confidence intervals is obtained for variances based on the number of experiments. An algorithm is given for finding the required number of experiments when the variance of estimation errors is not specified. A method is developed for obtaining the distribution density of the relative estimation error, its average value, the spread, and the probability of falling within a certain range of values, with a given degree of accuracy. We consider the determination of the required number of experiments (the number of questions or test tasks) for obtaining reasonable estimates by students based on the test results, as the practical implementation of the proposed methods of expert estimation.</p>
      </abstract>
      <kwd-group>
        <kwd>Sampling</kwd>
        <kwd>Estimation Error</kwd>
        <kwd>Confidence Interval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The expert estimation method is used to solve difficult to formalize problems for which
classical optimization methods cannot be applied. To choose the best solution, in this
case, the methods of voting (simple, weighted) and detachment of committees and
coalitions are used. In the case of collective estimation, the task of coordinating the
individual opinions of experts emerges. For random estimates, the minimum sum of the
coefficients of variation or the minimum sum of centered estimates (we call it the
estimation error) may be used as an indicator of decision consistency. Not only specialists
in a certain field of knowledge may be considered experts, but testing systems,
intelligent agents, algorithms, diagnostic systems, measuring instruments, and neural
networks as well. The use of expert estimates in the case of distance learning is of special
*
interest. Testing of each trainee may be considered a collective estimation of their
knowledge concerning different domains (topics, sections, lectures, practical training,
and laboratory classes) of a certain discipline. In this case, the collective body of experts
represents the questions and tasks of a test that are considered experiments.</p>
      <p>The reference literature pays much attention to the methods for determining the
optimal combination of the confidence interval and the confidence probability [4, 12], the
use of Chebyshev inequality for making the confidence intervals of expert estimates
[5], making the robust (stable and independent of the type of distribution law) estimates
[9], and the determination of the degree of various factors influence upon the result
based on expert estimation [7]. A recent trend is the use of artificial intelligence
methods in expert estimation: the methods of subjective probabilities [8] and fuzzy intervals
[10]. The expert estimation models are used to determine the quality of the education
process [11].</p>
      <p>The objective of the article is to develop mathematical models for consistent expert
estimates and estimates of collective decision quality.
2</p>
      <p>Setting the Task of Expert Estimation
Let us assume that n independent experts estimate some object A. Let xi i  i, n
1 n
denote the estimate given by the ith expert and let x   xi denote the average
estin i1
mate given by all experts. Data on the expertise conducted by those n experts may be
considered as the results of n experiments. All possible estimates of object A form a
general population of values for some random variable X.</p>
      <p>Before experimenting, we will consider the sample units X1, X2 ,..., Xn to be
pairwise independent random variables that have the same distribution law as X has. Let us
assume that X has a normal distribution, mx and  x is its mathematical expectation and
mean square deviation, accordingly. The variable x is random, since it is determined
by the sample [1], has a normal distribution (as the sum of normally distributed
varia 2
bles), and also M x  m , D x  x .</p>
      <p>x   n</p>
      <p>Let us assume that x is an arbitrary value of a random variable X. Then the variable
y  x  x characterizes the degree of consistency of the experts' opinions. This variable
will be called an absolute error of estimation.</p>
      <p>We assume that the interval  xmin , xmax  (here, xmin and xmax are correspondingly
the minimum and maximum values of the random variable X) coincides with the
interval u1,u2   mx  3 x , mx  3 x  and m  3 x  0 . If the first interval is less than the
x
second one, we stretch it by the appropriate number of times, moving to the new xmin
and xmax .</p>
    </sec>
    <sec id="sec-2">
      <title>Now the distribution density of the random variable Y may be found.</title>
      <p>The absolute error of estimation
Consider a random variable Y  X  x . Using the composition of two normally
distributed random variables, we may obtain the distribution law Y. When subtracting random
variables, their mathematical expectations are subtracted, and the variance is calculated
using the formula:
 n 1
 n
DY   D  X  x  D  X   D  x  2K X ,x  
    n 1
 n</p>
      <p>Dx , если x x1,..., xn,
Dx , если x x1,..., xn.</p>
      <p>The distribution density of the random variable Y is equal to:
 xx2 n
 n

 2  n 1  x
f  y   f  x  x  
 n xx2 n
 2  n 1  x
Note that
n  n 1 with an accuracy of 3.3 % if n  n0  40 .</p>
      <p>
e 2n1 x2 , если x x1,..., xn

e 2(n1) x2 , если x x1,..., xn
.</p>
      <p>
        (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
In this formula, my  0 and 99.73 % of Y values fall within the interval u1, u2  .
      </p>
      <p>
        It is inconvenient to make calculations using the formula (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) because  x2 is
unknown. Consider replacing  x2 with Sx2 
1 n 2
n 1 i1  xi  x . With such a replacement,
there will be a calculation error.
      </p>
      <p>Let us find the number n, at which the indicated replacement will take place with
the specified accuracy  and probability  . The problem is reduced to constructing a
confidence interval for the variance.</p>
      <p>Without violating generality, we compare two methods (described in [1] and [5])
for constructing a confidence interval for the variance.</p>
      <p>[1] discusses the approximate method for constructing a confidence interval for the
variance  x2 when the number of observations n  30 . It is the interval
 Sx2   t  Sx2   , Sx2   t  Sx2    , where  is the accuracy of
2 2 
 n 1 n 1 
the estimate and t is the argument of the Laplace function for the confidence
probability β. This suggests
n 
2t2  Sx2   </p>
      <p>2
 1.</p>
      <p>
        (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
[3] discusses the method for constructing a confidence interval for the variance when
n  30 . The confidence interval looks as follows:
      </p>
    </sec>
    <sec id="sec-3">
      <title>Then In this case,</title>
      <p> n 1 Sx2
 n 1  t

2(n 1)
,</p>
      <p>n 1 Sx2
n 1  t

.
2(n 1) 
</p>
      <p>
         2
n  8t2 Sx4  4  t2  0, 52  t4 1.
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
We obtain a sufficient condition for the second method to prevail over the first one
relative to the number of experts.
      </p>
      <p>
        Let us assume that   1Sx2 , i.e.  1 is the fraction of  in Sx2 . Then (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
will correspondingly look as follows:
n 
2t2  4t21  2t212 1,
      </p>
      <p>12
8t2  4t212  0, 52  t412 1.
n </p>
      <p>12
Now we find the difference  between the right-hand members of the latter inequalities
(subtracting the first one from the second one). The result is:
 
6t2  0, 52  t412  2t212 .</p>
      <p>12
Hence it is not difficult to demonstrate that   0 if and only if the following inequality
is correct:</p>
      <sec id="sec-3-1">
        <title>Maximum value t2  25 . Then   0 if and only if 1112  41  6  0.</title>
        <p>1 n
x   xi . For this purpose, estimates of n2 independent experts are considered.</p>
        <p>
          n i1
3. If the right member of (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) obtained in step 2 is not less than n2 and not more than
nmax , we get an estimate for the number of experts n so that the inequality: n2  n  nmax
is valid. Go to the End.
4. If the right member of (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) is less than n2, then we assume n2  n2 1. At that, if
n2  nmax , move to step 2.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Otherwise, it is necessary to decrease the value  or increase the value β.</title>
      <p>
        If n  n2 maxn0 , n1 the formula (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) is converted to the formula
since for these values of n  x2  Sx2 with probability 1 and with accuracy  .
      </p>
      <p>
        It may be demonstrated that the relative error of formula (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) does not exceed
f1  y   f1  x  x 
      </p>
      <p>1
2  Sx
xx2
e 2Sx2 ,
 3 
2 x
1 1</p>
      <p>(x  x)2   2 .
 x  n 1 1 2</p>
    </sec>
    <sec id="sec-5">
      <title>Thus,</title>
      <p>with probability  .
4</p>
      <p>
        The Relative Error of Estimation
When estimating the errors made by experts, an important role is assigned to the relative
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
(
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
error of estimation; such an error may be defined either as  
or by 1 
(the case x  x
x
is symmetric to  1 ).
      </p>
      <p>x  x
x
x  x
x
x  x
x
x  x
x</p>
    </sec>
    <sec id="sec-6">
      <title>Let us assume that  </title>
      <p>is a relative error. If
 1 , then the estimate x
is not consistent with x ; if</p>
      <p> 1, then x is consistent with x and the better the
smaller the relative error is.</p>
      <p>If  is no more than 3 %, then the estimate x has increased accuracy; if  falls
within the range from 3 % to 10 %, then this is the usual accuracy;  from 10 % to
20 % results in an approximate estimate [2].</p>
      <sec id="sec-6-1">
        <title>Consider 1 </title>
        <p>x  х
х</p>
        <p>as a relative error. This is the value of a random variable ∆1,
The relative error of estimate m :</p>
        <p>
          x
where  4  x  5 and 0   5  1 , i.e.
x  x
x
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
(
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
(
          <xref ref-type="bibr" rid="ref10">10</xref>
          )
where t ,n1 is the argument of the Student function  t, k  , which is such that
 t ,n1, n 1    1 and
[6] proved that  has a normal distribution if n  n3 and  2Sx2    max  2Sx2 , x 5,
        </p>
        <p>S S
x  nx t ,n1  mx  x  nx t ,n1 ,</p>
        <p> 2 2
n  n3  max n2 , Sx t ,n2 1  .</p>
        <p>  42 
 6 
x  mx   4 
mx mx</p>
        <p> 4 ,
x  4
represented by the ratio of variables Y  X  х and х 
distribution, moreover, M Y   0 , M х  m , if n  n2 with probability 1  and
x
with accuracy
 D  х  Sx2 .</p>
        <p>n   n</p>
        <p>
          One may replace an unknown value mx with an exact estimate x that meets the
condition:
while
where
Then
So, formulas (
          <xref ref-type="bibr" rid="ref11">11</xref>
          ) to (13) give a fairly accurate value of the distribution density of the
relative estimation error, its average value, the spread, and the probability of falling
within the range d, d  .
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Here is the algorithm for solving these problems.</title>
      <p>1. First, Algorithm 1 is applied. If it ends with Go to the End, move to step 2.</p>
    </sec>
    <sec id="sec-8">
      <title>2. n3 is calculated.</title>
      <sec id="sec-8-1">
        <title>3. If n2 is equal to n3, then n  n2 . Move to 4.</title>
        <p>
          Otherwise, we assume n2  n2 1 and move to step 2 of Algorithm 1.
4. Calculation of the spread of the relative estimation error.
4.1. First, we calculate “c” using the formula (
          <xref ref-type="bibr" rid="ref12">12</xref>
          ).
4.2. Use the right member of (
          <xref ref-type="bibr" rid="ref12">12</xref>
          ) for n  n1 .
5. Calculation of probability of the relative error falling within the interval d, d 
where d is specified.
        </p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>5.1. “c” is calculated using the formula (12).</title>
    </sec>
    <sec id="sec-10">
      <title>5.2. The value of the Laplace function at a point is calculated.</title>
      <p>d
Sxc
5</p>
      <p>Practical Implementation
When we tested students in the probability theory and mathematical statistics at the
Tver State Agriculture Academy, the maximum score was 10 points, taking into
account the complexity of the test. We used prompts (no more than 5); each of the prompts
reduces the score by 0.7 k points, where k is the number of useful prompts. The
maximum number of prompts, which is equal to 5, was determined based on these trial tests
that allowed consulting. The decrease in the score occurs in arithmetic progression. The
parameter 0.7 is selected from the condition of the maximum “penalty” for a prompt
because in this case the score is reduced for five prompts to the maximum level. If the
answer to this task is incorrect, two approaches are possible. The first (standard)
approach: a score of 0 points is graded for the task regardless of the number of prompts
the student used. The second approach is to use estimates of ontologies, various
fragments of this task, and methods of fuzzy control in an adequate system for estimating
the quality of teaching.</p>
      <p>When selecting the volume of n=300 test tasks, on average, the group of students
got: x  5,1; Sx2  3, 06 and  ≥ 292 if   0, 27 .</p>
      <p>When statistical data were approximated by the normal distribution law, the
significance level was 0.1.</p>
      <p>Note that the results of this article are valid for traditional testing without prompts
as well. At the same time, well-defined prompts greatly contribute to the use of the test
not only for monitoring but also for training, i.e. they increase the teaching potential of
test tasks.
6</p>
      <p>Conclusion
Creating reliable and high-quality methods of making collective decisions by experts is
a significant top-of-the-agenda topic of modern research in the field of complex systems
modeling. This issue is crucial for remote testing of learners.</p>
      <p>The developed method enables to obtain quantitative estimates of the required
number of experts to make a joint decision of a given quality.</p>
      <p>This method may be used not only for testing learners, but also in diagnostic
systems, product quality control, and other areas as well.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Ventcel',
          <string-name>
            <surname>E.S.</surname>
          </string-name>
          :
          <source>Probability Theory: Textbook. JUSTICE</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ganicheva</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          :
          <article-title>Mathematical Models and Methods for Evaluating Events</article-title>
          , Situations, and Processes: Textbook. Saint Petersburg: «Lan'» (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mihin</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          : Mathematical Statistics: Textbook. Moscow: MIREA (
          <year>2016</year>
          ).
          <article-title>(</article-title>
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Simankov</surname>
            ,
            <given-names>V.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchackaya</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teplouhov</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          :
          <article-title>Determination of the Optimal Combination of the Confidence Interval</article-title>
          and
          <string-name>
            <given-names>Confidence</given-names>
            <surname>Probability</surname>
          </string-name>
          .
          <source>The Bulletin of the Adyghe State University. Series 4: Natural-Mathematical and Technical Sciences</source>
          .
          <volume>3</volume>
          (
          <issue>246</issue>
          ),
          <fpage>69</fpage>
          -
          <lpage>74</lpage>
          (
          <year>2019</year>
          ). (In Russian)
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Suvorova</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pashchenko</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          , Tulup'eva, T.V.,
          <article-title>Tulup'ev,</article-title>
          <string-name>
            <surname>A.L.</surname>
          </string-name>
          :
          <article-title>Construction of Confidence Intervals for Assessing the Intensity of Risky Behavior Based on Chebyshev's Inequality</article-title>
          .
          <source>SPIIRAS Proceedings</source>
          .
          <volume>10</volume>
          ,
          <fpage>96</fpage>
          -
          <lpage>109</lpage>
          (
          <year>2009</year>
          ). DOI:
          <volume>10</volume>
          .15622/sp.10.6.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ganicheva</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          :
          <article-title>Test Technologies in Training. Tver: TGSHA (</article-title>
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Halikova</surname>
            ,
            <given-names>K.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ryzhkova</surname>
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Assessment of the Influence of Factors Based on Cognitive Modeling and Expert Assessment</article-title>
          .
          <source>Humanities Research</source>
          .
          <volume>2</volume>
          (
          <issue>54</issue>
          ),
          <fpage>300</fpage>
          -
          <lpage>303</lpage>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Boiko</surname>
          </string-name>
          , Ye.:
          <article-title>Methods of Forming an Expert Assessment of the Criteria of an Information System for Managing Projects and Programs Technology Transfer</article-title>
          .
          <source>Fundamental Principles and Innovative Technical Solutions. 2</source>
          ,
          <fpage>9</fpage>
          -
          <lpage>11</lpage>
          (
          <year>2018</year>
          ). DOI: http://dx.doi.org/10.21303/
          <fpage>2585</fpage>
          -
          <lpage>6847</lpage>
          .
          <year>2018</year>
          .
          <volume>00766</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Bonett</surname>
            ,
            <given-names>D.G.</given-names>
          </string-name>
          :
          <article-title>Robust Confidence Interval for a Ratio of Standard Deviations</article-title>
          . Applied Psychological Measurement.
          <volume>30</volume>
          (
          <issue>5</issue>
          ),
          <fpage>432</fpage>
          -
          <lpage>439</lpage>
          (
          <year>2006</year>
          ). DOI:
          <volume>10</volume>
          .1177/0146621605279551.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Structural Quantization of Vagueness in Linguistic Expert Opinions in an Evaluation Program</article-title>
          .
          <source>Fuzzy Sets and Systems</source>
          .
          <volume>119</volume>
          (
          <issue>1</issue>
          ),
          <fpage>171</fpage>
          -
          <lpage>186</lpage>
          (
          <year>2001</year>
          ) DOI:
          <fpage>10</fpage>
          .3745/KIPSTB.
          <year>2005</year>
          .
          <year>12B</year>
          .
          <volume>2</volume>
          .203.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Safargaliev</surname>
            ,
            <given-names>E.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eremina</surname>
            ,
            <given-names>I.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savitsky</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camelina</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          :
          <article-title>Mathematical Model and Qualimetric Assessment of Graduate Education Quality in Environment Saturated with Information and Communication Technologies</article-title>
          .
          <source>International Education Studies</source>
          .
          <volume>8</volume>
          (
          <issue>2</issue>
          ),
          <fpage>78</fpage>
          -
          <lpage>83</lpage>
          (
          <year>2015</year>
          ). DOI:
          <volume>10</volume>
          .5539/ies.v8n2p78.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Thulin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>On Split Sample and Randomized Confidence Intervals for Binomial Proportions</article-title>
          .
          <source>Statistics &amp; Probability Letters</source>
          .
          <volume>92</volume>
          ,
          <fpage>65</fpage>
          -
          <lpage>71</lpage>
          (
          <year>2014</year>
          ). DOI:
          <volume>10</volume>
          .1016/j.spl.
          <year>2014</year>
          .
          <volume>05</volume>
          .005.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>