=Paper= {{Paper |id=Vol-1297/104-123_paper-17 |storemode=property |title=Введение в анализ методов и средств поддержки научных экспериментов, движимых гипотезами (Introduction into Analysis of Methods and Tools for Hypothesis-Driven Scientific Experiment Support) |pdfUrl=https://ceur-ws.org/Vol-1297/104-123_paper-17.pdf |volume=Vol-1297 |dblpUrl=https://dblp.org/rec/conf/rcdl/KalinichenkoKKM14 }} ==Введение в анализ методов и средств поддержки научных экспериментов, движимых гипотезами (Introduction into Analysis of Methods and Tools for Hypothesis-Driven Scientific Experiment Support) == https://ceur-ws.org/Vol-1297/104-123_paper-17.pdf
         Introduction into Analysis of Methods and Tools
       for Hypothesis-Driven Scientific Experiment Support

© Kalinichenko L.A.              © Kovalev D.Y.                         © Kovaleva D.A.        © Malkov O.Y.
    Institute of Informatics Problems RAS                                   Institute of Astronomy RAS
                                           Moscow
leonidk@synth.ipi.ac.ru       dm.kovalev@gmail.com                       dana@inasan.ru              malkov@inasan.ru


                      Abstract                                  prerequisite for discoveries in the 21st century. Data
                                                                Intensive Research (DIR) denotes a crosscut of DIS/IT
    Data intensive sciences (DIS) are being                     areas aimed at the creation of effective data analysis
    developed in frame of the new paradigm of                   technologies for DIS and other data intensive domains.
    scientific study known as the Fourth paradigm,
                                                                    Science endeavors to give a meaningful description
    emphasizing       an    increasing    role   of
                                                                of the world of natural phenomena using what are
    observational, experimental and computer
                                                                known as laws, hypotheses and theories. Hypotheses,
    simulated data practically in all fields of
                                                                theories and laws in their essence have the same
    scientific study. The principal goal of data
                                                                fundamental character (Fig. 1) [48].
    intensive research (DIR) is an extraction
    (inference) of knowledge from data.                                             LAW            THEORY
    The intention of this work is to make an                                  Which might                  Which might
    overview of the existing approaches, methods                              become a                     become a
    and infrastructures of the data analysis in DIR                                                         An Explanatory
                                                                 An Generalizing       A Prediction
    accentuating the role of hypotheses in such                                                             Hypothesis
    research process and efficient support of                    Hypothesis
    hypothesis formation, evaluation and selection                                           The Word
    in course of the natural phenomena modeling
                                                                                            “Hypothesis”
    and experiments carrying out. An introduction
                                                                                            Could Mean
    into various concepts, methods and tools
    intended for effective organization of                               Fig. 1. Multiple incarnations of hypotheses
    hypothesis driven experiments in DIR is
                                                                    A scientific hypothesis is a proposed explanation of
    presented in the paper.
                                                                a phenomenon which still has to be rigorously tested. In
1 Hypotheses, theories, models and laws in                      contrast, a scientific theory has undergone extensive
                                                                testing and is generally accepted to be the accurate
data intensive science                                          explanation behind an observation. A scientific law is a
    Data intensive science (DIS) is being developed in          proposition, which points out any such orderliness or
accordance with the 4th Paradigm [29] of scientific             regularity in nature, the prevalence of an invariable
study (following three previous historical paradigms of         association between a particular set of conditions and
the science development (empirical science, theoretical         particular phenomena. In the exact sciences laws can
science, computational science)) emphasizing that               often be expressed in the form of mathematical
science as a whole is becoming increasingly dependent           relationships. Hypotheses explain laws, and well-tested,
on data as the core source for discovery. Emerging of           corroborated hypotheses become theories (Fig. 1). At
the 4th Paradigm is motivated by the huge amounts of            the same time the laws do not cease to be laws, just
data coming from scientific instruments, sensors,               because they did not appear first as hypotheses and pass
simulations, as well as from people accumulating data           through the stage of theories.
in Web or social nets. The basic objective of DIS is to             Though theories and laws are different kinds of
infer knowledge from the integrated data organized in           knowledge, actually they represent different forms of
networked infrastructures (such as warehouses, grids,           the    same      knowledge      construct.     Laws      are
clouds). At the same time, “Big Data” movement has              generalizations, principles or patterns in nature, and
emerged as a recognition of the increased significance          theories are the explanations of those generalizations.
of massive data in various domains. Open access to              However, classification expressed at the Fig. 1 is
large volumes of data therefore becomes a key                   subjective. [40] provides examples showing that the
                                                                differences between laws, hypotheses and theories
Proceedings of the 16th All-Russian Conference "Digital         consist only in that they stand at different levels in their
Libraries: Advanced Methods and Technologies, Digital           claim for acceptance depending on how much empirical
Collections" ― RCDL-2014, Dubna, Russia, October                evidence is amassed. Therefore there is no essential
13–16, 2014.                                                    difference between constructs used for expressing



                                                          104
hypotheses, theories and laws. Important role of                         In [52] the useful hypotheses of science are
hypotheses in scientific research can scarcely be                    considered to be of two kinds:
overestimated. In the edition of M. Poincaré's book [52]                 1. The          hypotheses          which          are
it is stressed that without hypotheses there is no science.          valuable precisely because they are either verifiable or
Thus it is not surprising that so much attention in the              else refutable through a definite appeal to the tests
scientific research and the respective publications is               furnished by experience;
devoted to the methods for hypothesis manipulation in
                                                                         2. The hypotheses which, despite the fact that
experimenting and modeling of various phenomena
                                                                     experience suggests them, are valuable despite, or
applying the means of informatics. The idea that the
                                                                     even because,     of     the    fact   that    experience
new approaches are needed that can address both data
                                                                     can neither confirm nor refute them.
driven and hypothesis driven sciences runs all through
this paper.       Such symbiosis alongside with the                      Aspects of science which are determined by the use
hypothesis-driven tradition of science (‘‘first                      of the hypotheses of the second kind are considered in
hypothesize-then-experiment’’) might cause wide                      the M. Poincaré's book [52] as “constituting an essential
application of another one that is typified by ‘‘first               human way of viewing nature, an interpretation rather
experiment-then-hypothesize’’ mode of research. Often                than a portrayal or a prediction of the objective facts of
the “first experiment” ordering in DIS is motivated by               nature, an adjustment of our conceptions of things to the
the necessity of analysis of the existing massive data to            internal needs of our intelligence”. According to M.
generate a hypothesis.                                               Poincaré's discussion, the central problem of the logic
                                                                     of science becomes the problem of the relation between
                     Generalization                                  the two fundamentally distinct kinds of hypotheses, i.e.,
   Abduction                                                         between those which cannot be verified or refuted
                        (Law)
                                                                     through experience, and those which can be empirically
     Induction                              Deduction                tested.
                                                                         The analysis in this paper will be focused mostly on
                                                                     the modeling of hypotheses of the first kind, leaving
                    Evidence (Facts)                                 issues of analysis the relations between such two kinds
                                                                     of hypotheses to further study.
       Fig. 2 Enhanced knowledge production diagram
                                                                         The rest of the paper is organized as follows.
    In the course of our study paying attention to the               Section 2 discusses the basic concepts defining the role
issue of inductive and deductive reasoning in hypothesis             of hypotheses in the formation of scientific knowledge
driven sciences will be emphasized. On Fig. 2 such                   and the respective organization of the scientific
ways of knowledge production are shown [48].                         experiments. Approaches for hypothesis formulation,
“Generalization” here means any subset of hypotheses,                logical reasoning, hypothesis modeling and testing are
theories and laws and “Evidence” is any subset of all                briefly introduced in Section 3. In Section 4 a general
facts accumulated in a specific DIS.                                 overview of the basic facilities provided by informatics
    All researchers collect and interpret empirical                  for the hypothesis driven experimentation scenarios,
evidence through the process called induction. This is a             including conceptual modeling, simulations, statistics
technique by which individual pieces of evidence are                 and machine learning methods is given. Into Section 5
collected and examined until a law is discovered or a                several examples of organization of hypothesis driven
theory is invented. Frances Bacon first formalized                   scientific experiments are included. Conclusion
induction [4]. The method of (naïve) induction (Fig. 2)              summarizes the discussion.
he suggested is in part the principal way by which
humans traditionally have produced generalizations that              2 Role of hypotheses in scientific
permit predictions. The problem with induction is that it            experiments: basic principles
is both impossible to collect all observations pertaining
to a given situation in all time – past, present and future.             Normally, scientific hypotheses have the form of a
                                                                     mathematical model. Sometimes one can also formulate
    The formulation of a new law begins through
                                                                     them as existential statements, stating that some
induction as facts are heaped upon other relevant facts.
                                                                     particular instance of the phenomenon under
Deduction is useful in checking the validity of a law.
                                                                     examination has some characteristic and causal
The Fig. 2 shows that a valid law would permit the
                                                                     explanations, which have the general form of universal
accurate prediction of facts not yet known. Also an
                                                                     statements, stating that every instance of the
abduction [49] is the process of validating a given
                                                                     phenomenon has a particular characteristic (e.g., for
hypothesis      through      reasoning     by    successive
                                                                     all x, if x is a swan, then x is white). Scientific
approximation. Under this principle, an explanation is
                                                                     hypothesis considered as a declarative statement
valid if it is the best possible explanation of a set of
                                                                     identifies the predicted relationship (associative or
known data. Abductive validation is common practice
                                                                     causal) between two or more variables (independent and
in hypothesis formation in science. Hypothesis related
                                                                     dependent). In causal relationship a change caused by
logic reasoning issues are considered in more details in
                                                                     the independent variable is predicted in the dependent
section 3.
                                                                     variable. Variables are more commonly related in non-
                                                                     causal (associative) way [25].



                                                               105
    In experimental studies the researcher manipulates                falsification can be sudden and definitive. Einstein said:
the independent variable. The dependent variable is                   “No amount of experimentation can ever prove me
often referred to as consequence or the presumed effect               right; a single experiment can prove me wrong”. To
that varies with a change of the independent variable.                scientists and philosophers outside the Popperian belief
The dependent variable is not manipulated. It is                      [53], science operates mainly by induction
observed and assumed to vary with changes in the                      (confirmation), and also and less often by
independent variable. Predictions are made from the                   disconfirmation (falsification). Its language is almost
independent variable to the dependent variable. It is the             always one of induction. For this survey both
dependent variable that the researcher is interested in               philosophical treatment of hypotheses are acceptable.
understanding, explaining or predicting [25].                         Sometimes such way of reasoning is called the
    In case when a possible correlation or similar                    hypothetico-deductive method. According to it,
relation between variables is investigated (such as, e.g.,            scientific inquiry proceeds by formulating a hypothesis
whether a proposed medication is effective in treating a              in a form that could conceivably be falsified by a test on
disease, that is, at least to some extent and for some                observable data. A test that could and does run contrary
patients), a few cases in which the tested remedy shows               to predictions of the hypothesis is taken as a
no effect do not falsify the hypothesis. Instead,                     falsification of the hypothesis. A test that could but does
statistical tests are used to determine how likely it is that         not run contrary to the hypothesis corroborates the
the overall effect would be observed if no real relation              theory.
as hypothesized exists. If that likelihood is sufficiently                A scientific method involves experiment, to test the
small, the existence of a relation may be assumed. In                 ability of some hypothesis to adequately answer the
statistical hypothesis testing two hypotheses are                     question under investigation. A prediction enabled by
compared, which are called the null hypothesis and the                hypothesis suggests a test (observation or experiment)
alternative hypothesis. The null hypothesis states that               for the hypothesis thus becoming testable. If a
there is no relationship between the phenomena                        hypothesis does not generate any observational tests,
(variables) whose relation is under investigation, or at              there is nothing that a scientist can do with it.
least not of the form given by the alternative hypothesis.                For example, not testable hypothesis: "Our universe
The alternative hypothesis, as the name suggests, is the              is surrounded by another, larger universe, with which
alternative to the null hypothesis: it states that there is           we can have absolutely no contact"; not verifiable
some kind of relation.                                                (though testable) hypothesis: "There are other inhabited
    Alternative hypotheses are generally used more                    planets in the universe"; scientific hypothesis (both
often than null hypotheses because they are more                      testable and verifiable): "Any two objects dropped from
desirable to state the researcher’s expectations. But in              the same height above the surface of the earth will hit
any study that involves statistical analysis the                      the ground at the same time, as long as air resistance
underlying null hypothesis is usually assumed [25]. It is             is not a factor" (http://www.batesville.k12.in.us/physics/
important, that the conclusion “do not reject the null                phynet/aboutscience/hypotheses.html).
hypothesis” does not necessarily mean that the null                       A problem (research question) should be formulated
hypothesis is true. It suggests that there is not sufficient          as an issue of what relation exists between two or more
evidence against the null hypothesis in favor of the                  variables. The problem statement should be such as to
alternative hypothesis. Rejecting the null hypothesis                 imply possibilities of empirical testing otherwise this
suggests that the alternative hypothesis may be true.                 will not be a scientific problem. Problems and
    Any useful hypothesis will enable predictions by                  hypotheses being generalized relational statements
reasoning (including deductive reasoning). It might                   enable to deduce specific empirical manifestations
predict the outcome of an experiment in a laboratory                  implied by the problem and hypotheses. In this process
setting or the observation of a phenomenon in nature.                 hypotheses can be deduced from theory and from other
The prediction may also invoke statistics assuming that               hypotheses. A problem cannot be scientifically solved
a hypothesis must be falsifiable [53], and that one                   unless it is reduced to hypothesis form, because a
cannot regard a proposition or theory as scientific if it             problem is not directly testable [37].
does not admit the possibility of being shown false. The                  Most formal hypotheses connect concepts by
way to demarcate between hypotheses is to call                        specifying the expected relationships between
scientific those for which we can specify (beforehand)                propositions. When a set of hypotheses are grouped
one or more potential falsifiers as the respective                    together they become a type of conceptual framework.
experiments. Falsification was supposed to proceed                    When a conceptual framework is complex and
deductively instead of inductively.                                   incorporates causality or explanation it is generally
    Other philosophers of science have rejected the                   referred to as a theory [28]. In general, hypotheses have
criterion of falsifiability or supplemented it with other             to reflect the multivariate complexity of the reality. A
criteria, such as verifiability (only statements about the            scientific theory summarizes a hypothesis or a group of
world that are empirically confirmable or logically                   hypotheses that have been supported with repeated
necessary are cognitively meaningful). They claim that                testing. A theory is valid as long as there is no evidence
science proceeds by "induction"— that is, by finding                  to dispute it. Scientific paradigm explains the working
confirming instances of a conjecture. Popper treated                  set of theories under which science operates.
confirmation as never certain [53]. However, a



                                                                106
    Elements of hypothesis-driven research and their
relationships are shown on Fig. 3 [23, 57]. The
hypothesis triangle relations, explains, formulates,
                                                                        Weak              Galaxy              Earth special location
represents are functional in the scientist's final decision in
                                                                        lensing           clustering
adopting a particular model m1 to formulate a hypothesis
h1, which is meant to explain phenomenon p1.
                                                                            Dark Energy                     Non uniform universe


                                                                          Fig. 4. A lattice theoretic representation for hypothesis
                                                                                                 relationship
                                                                            The hypothesis lattice is unfolded into model and
                                                                       phenomena isomorphic lattices according to the
                                                                       hypothesis triangle (Fig. 3) [23]. The lattices are
                                                                       isomorphic if one takes subsets of M (Model), H
                                                                       (Hypotheses) and P (Phenomenon) such that formulates,
                                                                       explains and represents are both one-to-one and onto
                                                                       mappings (i.e., bijections), seen as structure-preserving
        Fig. 3. Elements of hypothesis-driven research                 mappings (morphisms). Example of the isomorphic
    In [23] the lattice structure for hypothesis                       lattice is shown on the Fig. 5 [23]. This particular lattice
interconnection is proposed as shown on Fig. 4. A                      corresponds to the case in Computational
hypothesis lattice is formed by considering a set of                   Hemodynamics considered in [23]. Here model m1
hypotheses equipped with wasDerivedFrom as a strict                    formulates hypothesis h1, which explains phenomenon
order < (from the bottom to the top). Hypotheses                       p1. Similarly, m2 formulates h2, which explains p2, and
directly derived from exactly one hypothesis are atomic,               so on. Properties of the hypothesis lattices and
while those directly derived from at least two                         operations over them are considered in [24].
hypotheses are complex.




                                             h1. Mass     h2. Incompressible h3. Momentum h4. Viscous h5. No
                                             conversation fluid              conservation fluid       body forces

                                                                           h6. blend(h1,h2)     h7. blend(h3,h4,h5)



                                                                                      h8. blend(h6,h7)

          m1.            m2.           m3.               m4.                  m5.
                                                                                                 explains




                  m6.                m7.



                           m8.

                                                p1. Mass p2. Fluid   p3. Momentum p4. Fluid                    p5. Body
                                                net flux compression net flux     friction                     force effects

                                                                         p6. Fluid        p7. Fluid
                                                                         divergence       dynamics

                                                                                  p8. Fluid
                                                                                  behavior

                        Fig. 5. Hypothesis lattice unfolded into model and phenomenon isomorphic lattice




                                                                 107
    Models are one of the principal instruments of                 for phenomena, so additionally some criteria for
modern science. Models can perform two fundamentally               choosing among different hypotheses are required.
different representational functions: a model can be a                 One way to implement abductive model of reasoning
representation of a selected part of the world, or a model         is the abductive logic programming [36]. Hypothesis
can represent a theory in the sense that it interprets the         generation in abduction logical framework is organized
laws and hypotheses of that theory.                                as follows. During the experiment, some new
    Here we consider scientific models to be                       observations are encountered. Let B represents the
representations in both senses at the same time. One of            background knowledge; O is the set of facts that
the most perplexing questions in connection with                   represents observations. Both B and O are logic
models is how they relate to theories. In this respect             programs (set of rules in some rule language). In
models can be considered as a complement to theories,              addition, Γ stands for a set of literals representing the set
as preliminary theories, can be used as substitutions of           of abducibles, which are candidate assumptions to be
theories when the latter are too complicated to handle.            added to B for explaining O. Given B, O and Γ, the
Learning about the model is done through experiments,              hypothesis-generation problem is to find a set H of
thought experiments and simulation. Given a set of                 literals (called a hypothesis) such that: 1) B and H entail
parameters, a model can generate expectations about                O, 2) B and H is consistent, and 3) H is some subset of
how the system will behave in a particular situation. A            Г. If all conditions are met then H is an explanation of
model and the hypotheses it is based upon are supported            O (with respect to B and Γ). Examples of abductive
when the model generates expectations that match the               logic programming systems include ACLP [35], A-
behavior of its real-world counterpart.                            system [71], ABDUAL [2] and ProLogICA [59].
                                                                   Abductive logic programming can also be implemented
    A law generalizes a body of observations. Generally,           by means of Answer Set Programming systems, e.g. by
a law represents a group of related undisputable                   the DLV system [14].
hypotheses using a handful of fundamental concepts and                 The example abductive logic program in ProLogICA
equations to define the rules governing a set of                   describes a simple model of the lactose metabolism of
phenomena. A law does not attempt to explain why                   the bacterium E.Coli [59]. The background knowledge
something happens – it simply states that it does.                 B describes that E. coli can feed on the sugar lactose if it
    Facilities for support of the hypothesis-driven                makes two enzymes permease and galactosidase. Like
experimentation will be discussed in the remaining                 all enzymes (E), these are made if they are coded by a
sections.                                                          gene (G) that is expressed. These enzymes are coded by
                                                                   two genes (lac(y) and lac(z)) in cluster of genes (lac(X))
3 Hypothesis manipulation in scientific                            called an operon that is expressed when the amounts
                                                                   (amt) of glucose are low and lactose are high or when
experiments                                                        they are both at medium level. The abducibles, Г,
                                                                   declare all ground instances of the predicates "amount"
3.1 Hypothesis generation                                          as assumable. This reflects the fact that in the model it is
Researchers that support rationality of scientific                 not known what are the amounts at any time of the
discovery presented several methods for hypothesis                 various substances. This is incomplete information that
generation, including discovery as abduction, induction,           we want to find out in each problem case that we are
anomaly detection, heuristics programming and use of               examining. The integrity constraints state that the
analogies [73].                                                    amount of a substance (S) can only take one value.
    Discovery as abduction characterizes reasoning                 ## Background Knowledge (B)
                                                                   feed(lactose):-
processes that take place before a new hypothesis is               make(permease),make(galactosidase).
justified. The abductive model of reasoning that leads to          make(Enzyme):- code(Gene,Enzyme),express(Gene).
plausible hypotheses formulation is conceptualized as              express(lac(X)):-
an inference beginning with data. According to [50] an             amount(glucose,low),amount(lactose,hi).
                                                                   express(lac(X)):-
abduction happens as follows: 1) Some phenomena p1,                amount(glucose,medium),amount(lactose,medium).
p2, p3, … are encountered for which there is no or little          code(lac(y),permease).
explanation; 2) However, p1, p2, p3, … would not be                code(lac(z),galactosidase).
surprising if a hypothesis H were added. They would                temperature(low):-amount(glucose,low).
                                                                   false :- amount(S,V1), amount(S,V2), V1 != V2.
certainly follow from something like H and would be
explained by it; 3) Therefore there is good reason for             ## Abducibles (Г)
elaborating an hypothesis H – for proposing it as a                abducible_predicate(amount).
possible hypothesis from which the assumption p1, p2,              ## Observation (O)
p3, … might follow. The abductive model of reasoning               feed(lactose).
is primarily a process of explaining anomalies or
surprising phenomena [63]. The scientists' reasoning               This goal generates two possible hypotheses:
                                                                   {amount(lactose,hi), amount(glucose,low)}
proceeds abductively from an anomaly to an                         {amount(lactose,medium),amount(glucose,medium)}
explanatory hypothesis in light of which the phenomena
would no longer be surprising. There can be several                    Just a couple of another examples of real rule-based
different hypotheses that can serve as the explanations            systems, where abductive logic programming is used.




                                                             108
Robot Scientist (see 4.4) abductively hypothesizes new              reasoning with hypotheses, i.e., the propositions whose
facts about the yeast functional biology by inferring               truth or falsity is uncertain.
what is missing from a model [38]. In [68], both                        Bayesian probability belongs to the category of
abduction and induction are used to formulate                       evidential probabilities; to evaluate the probability of a
hypotheses about inhibition in metabolic pathways.                  hypothesis, the Bayesian probabilist specifies some
Augmenting background knowledge is done with                        prior probability, which is then updated in the light of
abduction, after that induction is used for learning                new, relevant data (evidence) [64]. The Bayesian
general rules. In [33] authors use SOLAR reasoning                  interpretation provides a standard set of procedures and
system to abductively generate hypotheses about the                 formulae to perform this calculation.
inhibitory effects of toxins on the rat metabolisms.
                                                                        Hypothesis testing in classical statistic style. After
    The process of discovery is deeply connected also
                                                                    null and alternative hypotheses are stated, some
with the search of anomalies. There are a lot of methods
                                                                    statistical assumptions about data samples should be
and algorithms to discover anomalies. Anomaly
                                                                    done, e.g. assumptions about statistical independence or
detection is an important research problem in data
                                                                    distributions of observations. Failing to provide correct
mining that aims to find objects that are considerably
                                                                    assumptions leads to the invalid test results.
dissimilar, exceptional and inconsistent with respect to
the majority data in an input database [6].                             A common problem in classical statistics is to ask
                                                                    whether a given sample is consistent with some
    Analogies play several roles in science. Not only do
                                                                    hypothesis. For example, we might be interested in
they contribute to discovery but they also play a role in
                                                                    whether a measured value xi, or the whole set {xi}, is
the development and evaluation of scientific theories
                                                                    consistent with being drawn from a Gaussian
(new hypotheses) by analogical reasoning.
                                                                    distribution N(μ,σ). Here N(μ,σ) is our null hypothesis.
3.2 Hypothesis evaluation                                               It is always assumed that we know how to compute
                                                                    the probability of a given outcome from the null
    Being testable and falsifiable, a scientific hypothesis         hypothesis: for example, given the cumulative
provides a solid basis to its further modeling and testing.         distribution function, 0 ≤ H0(x) ≤ 1, the probability that
There are several ways to do it, including the use of               we would get a value at least as large as xi is
statistics, machine learning and logic reasoning                    p(x > xi ) = 1 – H0(xi), and is called the p-value.
techniques.                                                         Typically, a threshold p value is adopted, called the
3.2.1 Statistical testing of hypotheses                             significance level α, and the null hypothesis is rejected
    The classical (frequentist) and Bayesian statistic              when p ≤  (e.g., if  = 0.05 and p < 0.05, the null
approaches are applicable for hypothesis testing and                hypothesis is rejected at a 0.05 significance level). If we
selection. Brief summary of the basic differences                   fail to reject a hypothesis, it does not mean that we
between these approaches are as follows [34].                       proved its correctness because it may be that our sample
                                                                    is simply not large enough to detect an effect.
    Classical (frequentist) statistics is based on the
following beliefs:                                                      When performing these tests, we can meet with two
                                                                    types of errors, which statisticians call Type I and Type
    − Probabilities refer to relative frequencies of                II errors. Type I errors are cases when the null
events. They are objective properties of the real world;            hypothesis is true but incorrectly rejected. In the context
    − Parameters of hypotheses (models) are fixed,                  of source detection, these errors represent spurious
unknown constants. Because they are not fluctuating,                sources, or more generally, false positives (with respect
probability      statements    about    parameters     are          to the alternative hypothesis). The false-positive
meaningless;                                                        probability when testing a single datum is limited by the
    − Statistical procedures should have well-defined               adopted significance level . Cases when the null
long-run frequency properties.                                      hypothesis is false, but it is not rejected are called
    In contrast, Bayesian approach takes the following              Type II errors (missed sources, or false negatives (again,
assumptions:                                                        with respect to the alternative hypothesis)). The false-
                                                                    negative probability when testing a single datum is
    − Probability describes the degree of subjective                usually called , and is related to the power of  test as
belief, not the limiting frequency. Probability statements
                                                                    (1 – ). Hypothesis testing is intimately related to
can be made about things other than data, including
                                                                    comparisons of distributions.
hypotheses (models) themselves as well as their
parameters;                                                             As the significance level  is decreased (the
                                                                    criterion for rejecting the null hypothesis becomes more
    − Inferences about a parameter are made by
                                                                    conservative), the number of false positives decreases
producing its probability distribution — this distribution
                                                                    and the number of false negatives increases. Therefore,
quantifies the uncertainty of our knowledge about that
                                                                    there is a trade-off to be made to find an optimal value
parameter. Various point estimates, such as expectation
value, may then be readily extracted from this                      of , which depends on the relative importance of false
distribution.                                                       negatives and positives in a particular problem. Both the
                                                                    acceptance of false hypotheses and the rejection of true
    The Bayesian interpretation of probability can be               ones are errors that scientists should try to avoid. There
seen as an extension of propositional logic that enables            is discussion as to what states of affairs is less desirable;



                                                              109
many people think that the acceptance of a false                             The question often arises as to which is the ‘best’
hypothesis is always worse than failure to accept a true                model (hypothesis) to use; ‘model selection’ is a
one and that science should in the first place try to avoid             technique that can be used when we wish to
the former kind of error.                                               discriminate between competing models (hypotheses)
    When many instances of hypothesis testing are                       and identify the best model (hypothesis) in a set, {M1,
performed, a process called multiple hypothesis testing, the            ..., Mn}, given the data.
fraction of false positives can significantly exceed the                     We need to remind the basic notation. The Bayes
value of . The fraction of false positives depends not only            theorem can be applied to calculate the posterior
on  and the number of data points, but also on the number              probability p(Mj|d) for each model (or hypothesis) Mj
of true positives (the latter is proportional to the number of          representing our state of knowledge about the truth of
instances when an alternative hypothesis is true).                      the model (hypothesis) in the light of the data d as
    Depending on data type (discrete vs. continuous                     follows:
random variables) and what we can assume (or not)                                      p(Mj|d) = p(d|Mj) p(Mj) / p(d),
about the underlying distributions, and the specific                    where p(Mj) is the prior belief in the model (hypothesis)
question we ask, we can use different statistical tests.                that represents our state of knowledge (or ignorance)
The underlying idea of statistical tests is to use data to              about the truth of the model (hypothesis) before we
compute an appropriate statistic, and then compare the                  have analyzed the current data, p(d|Mj) is the model
resulting data-based value to its expected distribution.                (hypothesis) likelihood (represents the probability that
The expected distribution is evaluated by assuming that                 some data are produced under the assumption of this
the null hypothesis is true. When this expected                         model) and p(d) is a normalization constant given by:

                                                                                                 p (d M ) p ( M ) .
distribution implies that the data-based value is unlikely
to have arisen from it by chance (i.e., the corresponding                              p(d) =              i         i
                                                                                                   i
p value is small), the null hypothesis is rejected with                     The relative ‘goodness’ of models is given by a
some threshold probability , typically 0.05 or 0.01                    comparison of their posterior probabilities, so to
(p < ). Note again that p >  does not mean that the                   compare two models Ma and Mb, we look at the ratio of
hypothesis is proven to be correct.                                     the model posterior probabilities:
    The number of various statistical tests in the                          p(Ma|d) / p(Mb|d) = p(d|Ma) p(Ma) / p(d|Mb) p(Mb).
literature is overwhelming and their applicability is
                                                                            The Bayes factor, Bab can be computed as the ratio
often hard to decide (see [19, 31] for variety of
                                                                        of the model likelihoods:
statistical methods in SPSS). When the distributions are
not known, tests are called nonparametric, or                                            Bab = p(d|Ma) / p(d|Mb).
distribution-free tests. The most popular nonparametric                     Empirical scale for evaluating the strength of
test is the Kolmogorov–Smirnov (K-S) test, which                        evidence from the Bayes factor Bij between two models
compares the cumulative distribution function, F (x), for               is shown in Tabl. 1 [45].
two samples, {x1i }, i = 1, ..., N1 and {x2i }, i = 1, ..., N2.
                                                                              Tabl. 1. Strength of evidence for Bayes factor Bij
The K-S test is not the only option for nonparametric                                          for two models
comparison of distributions. The Cramér-von Mises
criterion, the Watson test, and the Anderson–Darling                        |ln Bij|        Odds               Strength of evidence
test are similar in spirit to the K-S test, but consider                    < 1.0         <3:1          Inconclusive
somewhat different statistics. The Mann–Whitney–                             1.0          ~3:1          Weak evidence
Wilcoxon test (or the Wilcoxon rank-sum test) is a                           2.5         ~ 12 : 1       Moderate evidence
nonparametric test for testing whether two data sets are                     5.0         ~ 150 : 1      Strong evidence
drawn from distributions with different location
parameters (if these distributions are known to be                          The Bayes factor gives a measure of the ‘goodness’
Gaussian, the standard classical test is called the t test).            of a model, regardless of the prior belief about the
A few standard statistical tests can be used when we                    model; the higher the Bayes factor, the better the model
know, or can assume, that both h(x) and f(x) are                        is. In many cases, the prior belief in each model in the
Gaussian distributions (e.g., the Anderson–Darling test,                set of proposed models will be equal, so the Bayes
the Shapiro–Wilk test) [34]. More on statistical tests can              factor will be equivalent to the ratio of the posterior
be found in [19, 31, 32, 34].                                           probabilities of the models. The ‘best’ model in the
                                                                        Bayesian sense is the one which gives the best fit to the
    Hypothesis (model) selection and testing in                         data with the smallest parameter space.
Bayesian style. The Bayesian approach can be thought                        A special case of model (hypothesis) selection is
of as formalizing the process of continually refining our               Bayesian hypothesis testing [34, 62]. Taking M1 to be
state of knowledge about the world, beginning with no                   the “null” hypothesis, we can ask whether the data
data (as encoded by the prior), then updating that by                   supports the alternative hypothesis M2, i.e., whether we
multiplying in the likelihood once the data are observed                can reject the null hypothesis. Taking equal priors
to obtain the posterior. When more data are taken, then                 p(M1) = p(M2), the odds ratio is
the posterior based on the first data set can be used as
the prior for the second analysis. Indeed, the data sets                                 B21 = p(d|M1) / p(d|M2).
can be different.




                                                                  110
    The inability to reject M1 in the absence of an                 a good inductive argument the premises should provide
alternative hypothesis is very different from the                   some degree of support for the conclusion, where such
hypothesis testing procedure in classical statistics. The           support means that the truth of the premises indicates
latter procedure rejects the null hypothesis if it does not         with some degree of strength that the conclusion is true.
provide a good description of the data, that is, when it is         If the logic of good inductive arguments is to be of any
very unlikely that the given data could have been                   real value, the measure of support it articulates should
generated as prescribed by the null hypothesis. In                  meet the Criterion of Adequacy (CoA): as evidence
contrast, the Bayesian approach is based on the                     accumulates, the degree to which the collection of true
posterior rather than on the data likelihood, and cannot            evidence statements comes to support a hypothesis, as
reject a hypothesis if there are no alternative                     measured by the logic, should tend to indicate that the
explanations for observed data [34].                                hypotheses are probably false or probably true. In [27]
    Comparing classical and Bayesian approaches [34],               the extent to which a kind of logic based on the Bayes
it is rare for a mission-critical analysis be done in the           theorem can estimate how the implications of
“fully Bayesian” manner, i.e., without the use of the               hypotheses about evidence claims influences the degree
frequentist tools at the various stages. Philosophy and             to which hypotheses are supported is discussed in detail.
beauty aside, the reliability and efficiency of the                 In particular, it is shown how such a logic may be
underlying computations required by the Bayesian                    applied to satisfy the CoA: as evidence accumulates,
framework are the main practical issues. A central                  false hypotheses will very probably come to have
technical issue at the heart of this is that it is much             evidential support values (as measured by their
easier to do optimization (reliably and efficiently) in             posterior probabilities) that approach 0; and as this
high dimensions than it is to do integration in high                happens, a true hypothesis will very probably acquire
dimensions. Thus the usable machine learning methods,               evidential support values (measured by their posterior
while there are ongoing efforts to adapt them to                    probabilities) that approach 1.
Bayesian framework, are almost all rooted in frequentist            3.2.3 Parameter estimation
methods.
                                                                        Models (hypotheses) are typically described by
    Most users of Bayesian estimation methods, in                   parameters θ whose values are to be estimated from
practice, are likely to use a mix of Bayesian and                   data. We describe this process according to [34]. For a
frequentist tools. The reverse is also true—frequentist             particular model M and prior information I we get:
data analysts, even if they stay formally within the
frequentist framework, are often influenced by                               p(M, θ|d, I) = p(d|M, θ, I) p(M, θ|I) / p(d|I)
“Bayesian thinking,” referring to “priors” and                          The result p(M, θ|d, I) is called the posterior
“posteriors.” The most advisable position is probably to            probability density function (pdf) for model M and
know both paradigms well, in order to make informed                 parameters θ, given data d and other prior information I.
judgments about which tools to apply in which                       This term is a (k + 1)-dimensional pdf in the space
situations [34]. More details on Bayesian style of                  spanned by k model parameters and the model M. The
hypothesis testing can be found in [34, 62, 64].                    term p(d|M, θ, I) is the likelihood of data given some
3.2.2 Logic-based hypothesis testing                                model M and some fixed values of parameters θ
                                                                    describing it, and all other prior information I. The term
     According to the hypothetico-deductive approach
                                                                    p(M, θ|I) is the a priori joint probability for model M
the hypotheses are tested by deducing predictions or
                                                                    and its parameters θ in the absence of any of the data
other empirical consequences from general theories. If
                                                                    used to compute likelihood, and is often simply called
these predictions are verified by experiments, this
                                                                    the prior.
supports the hypothesis. It should be noted that not
anything that is logically entailed by a hypothesis can be              In the Bayesian formalism, p(M, θ|d, I) corresponds
confirmed by a proper test for it. The relation between             to the state of our knowledge (i.e., belief) about a model
hypothesis and evidence is often empirica l rather than             and its parameters, given data d. To simplify the
logical. A clean deduction of empirical consequences                notation, M(θ) will be substituted by M whenever the
from a hypothesis, as it may sometimes exist in physics,            absence of explicit dependence on θ is not confusing. A
is practically inapplicable in biology. Thus, entailment            completely Bayesian data analysis has the following
of the evidence by hypotheses under test is neither                 conceptual steps:
sufficient nor necessary for a good test. Inference to the              1. Formulation of the data likelihood p(d|M, I).
best explanation is usually construed as a form of                      2. Choice of the prior p(θ|M,I), which incorporates
inductive inference (see abduction in 3.1) where a                  all other knowledge that might exist, but is not used
hypothesis’ explanatory credentials are taken to indicate           when computing the likelihood (e.g., prior
its truth [72].                                                     measurements of the same type, different
     An inductive logic is a system of evidential support           measurements, or simply an uninformative prior).
that extends deductive logic to less-than-certain                   Several methods for constructing "objective" priors
inferences. For valid deductive arguments the                       have been proposed. One of them is the principle of
premises logically entail the conclusion, where the                 maximum entropy for assigning uninformative priors by
entailment means that the truth of the premises provides            maximizing the entropy over a suitable set of pdfs,
a guarantee of the truth of the conclusion. Similarly, in           finding the distribution that is least informative (given




                                                              111
the constraints). Entropy maximization with no testable            accuracy and properties of the algorithms (such as, e.g.,
information takes place under a single constraint: the             their convergence if they are iterative) are the issues to
sum of the probabilities must be one. Under this                   be investigated. Machine learning algorithms focus on
constraint, the maximum entropy for a discrete                     prediction, based on known properties learned from the
probability distribution is given by the uniform                   training data. Such machine learning algorithms as
distribution.                                                      decision tree, association rule, neural networks, support
    3. Determination of the posterior p(M|d, I), using             vector machines as well as other techniques of learning
Bayes theorem above. In practice, this step can be                 in Bayesian and probabilistic models [5, 26] are
computationally        intensive      for      complex             examples of the methods that belong to this second
multidimensional problems.                                         culture.
    4. The search for the best model M parameters,                     The models that best emulate the nature in terms of
which maximizes p(M|d, I), yielding the maximum a                  predictive accuracy are also the most complex and
posteriori (MAP) estimate. This point estimate is the              inscrutable. Nature forms the outputs y from the inputs x
natural analog to the maximum likelihood estimate                  by means of a black box with complex and unknown
(MLE) from classical statistics.                                   interior. Current accurate prediction methods are also
                                                                   complex black boxes (such as neural nets, forests,
    5. Quantification of uncertainty in parameter
                                                                   support vectors). So we are facing two black boxes,
estimates, via credible regions. As in MLE, such an
                                                                   where ours seems only slightly less inscrutable than
estimate can be obtained analytically by doing
                                                                   nature’s [10]. In a choice between accuracy and
mathematical derivations specific to the chosen model.
                                                                   interpretability, in applications people sometimes prefer
Also as in MLE, various numerical techniques can be
                                                                   interpretability.
used to simulate samples from the posterior. This can be
viewed as an analogy to the frequentist approach, which                However, the goal of a model is not interpretability
can simulate draws of samples from the true underlying             (a way of getting information), but getting useful,
distribution of the data. In both cases, various                   accurate information about the relation between the
descriptive statistics can then be computed on such                response and predictor variables. It is stated in [10] that
samples to examine the uncertainties surrounding the               algorithmic models can give better predictive accuracy
data and estimators of model parameters based on that              than formulaic models, providing also better
data.                                                              information about the underlying mechanism. And
                                                                   actually this is what the goal of statistical analysis is.
    6. Hypothesis testing as needed to make other
                                                                   The researchers should be focused on solving the
conclusions about the model (hypothesis) or parameter
                                                                   problems instead of asking what regression model they
estimates.
                                                                   can create.
3.3 Algorithmic generation and evaluation                              An objection to this idea (expressed by Cox) is that
of hypotheses                                                      prediction without some understanding of underlying
                                                                   process and linking with other sources of information
    Two cultures of data analysis (formulaic modeling 1            becomes more and more tentative. Due to that it is
and algorithmic modeling) distinguished here in                    suggested to construct the stochastic calculation models
accordance with [10] can be applied to the hypothesis              that summarize the understanding of the phenomena
extraction and generation based on data.                           under study. One of the objectives of such approach
    Formulaic modeling is a process for estimating the             might be an understanding and test of hypotheses about
relationships among variables. It includes many                    underlying process. Given the relatively small sample
techniques for modeling and analyzing several                      size following such direction could be productive. But
variables, when the focus is on the formulae y = f(x)              data characteristics are rapidly changing. In many of the
that give a relation specifying a vector of dependent              most interesting current problems, the idea of starting
variables y in terms of a vector of independent variables          with a formal model is not tenable. The methods used in
x. In a statistics experiment (based on various regression         statistics for small sample sizes and a small number of
techniques) the dependent variable defines the event               variables are not applicable. Data analytics need to be
studied and is expected to change whenever the                     more pragmatic. Given a statistical problem, find a good
independent variable (predictor variables, extraneous              solution, whether it is a formulaic model, an algorithmic
variables) is altered. Such methods as linear regression,          model or a Bayesian model or a completely different
logistic regression, multiple regression are well-known            approach.
examples of the representatives of this modeling                       In the context of the hypothesis driven analysis we
approach.                                                          should pay attention to the question how far can we go
    In the algorithmic modeling culture the approach is            applying the algorithmic modeling for hypothesis
to find an algorithm that operates on x to predict the             generation and testing. Various approaches to machine
responses y. What is observed is a set of x’s that go in           learning use related to hypothesis formation and
and a subsequent set of y’s that come out. Predictive              selection can be found in [5, 10, 34].
                                                                       Besides machine learning, an interesting example of
1
  In [10] instead of “formulaic modeling” the term “data           algorithmic generation of hypotheses can be found in
modeling” is used that looks misleading in the computer            the IBM Watson project [18] where the symbiosis of the
science context.                                                   general-purpose reusable natural language processing



                                                             112
(NLP) and knowledge representation and reasoning                   amount of observations nowadays available. The so-
(KRR) technologies (under the name DeepQA) is                      called ‘cosmological concordance model’ is based on
exploited for answering arbitrary questions over the               the cosmological principle (i.e. the Universe is isotropic
existing natural language documents as well as                     and homogeneous, at least on large enough scales) and
structured data resources. Hypothesis generation takes             on the hot big bang scenario, complemented by an
the results of question analysis and produces candidate            inflationary epoch. This remarkably simple model is
answers by searching the available data sources and                able to explain with only half a dozen free parameter
extracting answer-sized snippets from the search results.          observations spanning a huge range of time and length-
Each candidate answer plugged back into the question is            scales. Since both a cold dark matter (CDM) and a
considered a hypothesis, which the system has to prove             cosmological constant (Λ) component are required to fit
correct with some degree of confidence. After merging,             the data, the concordance model is often referred to as
the system must rank the hypotheses and estimate                   ‘the ΛCDM model’.
confidence based on their merged scores. A machine-                    Several different types of explanation are possible
learning approach adopted is based on running the                  for the apparent late time acceleration of the Universe,
system over a set of training questions with known                 including different classes of dark energy model such as
answers and training a model based on the scores. An               ΛCDM, wCDM; theories of modified gravity; void
important consideration in dealing with NLP-based                  models or the back reaction [45]. The methodology of
scorers is that the features they produce may be quite             Bayesian doubt which gives an absolute measure of the
sparse, and so accurate confidence estimation requires             degree of goodness of a model has been applied to the
the application of confidence-weighted learning                    issue of whether the ΛCDM model should be doubted.
techniques [18] – a new class of online learning
                                                                       The methodology of Bayesian doubt dictates that an
methods that maintain a probabilistic measure of
                                                                   unknown idealized model X should be introduced
confidence in each parameter. It is important to note that
                                                                   against which the other models may be compared.
instead of statistics based hypothesis testing, contextual
                                                                   Following [67], ‘doubt’ may be defined as the posterior
evaluation of a wide range of loosely coupled
                                                                   probability of the unknown model:
probabilistic question and semantic based content
analytics is applied for scoring different questions                             D ≡ p(X|d) = p(d|X) p(X) / p(d).
(hypotheses) and content interpretations. Training                     Here p(X) is the prior doubt, i.e. the prior on the
different models on different portions of the data in              unknown model, which represents the degree of belief
parallel and combining the learned classifiers into a              that the list of known models does not contain the true
single classifier allows to make the process applicable to         model. The sum of all the model priors must be unity.
the large collections of data. More details on that can be             The methodology of Bayesian doubt requires a
found in [17, 18] as well as in other Watson project               baseline model (the best model in the set of known
related publications.                                              models), for which in this application the ΛCDM has
                                                                   been chosen. The average Bayes factor between ΛCDM
3.4 Bayesian motivation for discovery
                                                                   and each of the known models is given by:
    One way for discriminating between competing
                                                                                                   B .
                                                                                                     N
models of some phenomenon is to use Bayesian model                                   ≡ 1/N             i
                                                                                                     i 1
selection approach (3.2.1), the Bayesian evidences for
                                                                       The ratio R between the posterior doubt and prior
each of the proposed models (hypotheses) can be
                                                                   doubt, which is called the relative change in doubt, is:
computed and the models can then be ranked by their
Bayesian evidence. This is a good method for                                               R ≡ D/p(X).
identifying which is the best model in a given set of                  For doubt to grow, i.e. the posterior doubt to be
models, but it gives no indication of the absolute                 greater than the prior doubt (R << 1), the Bayes factor
goodness of the model. Bayesian model selection says               between the unknown model X and the baseline model
nothing about the overall quality of the set of models             must be much greater than the average Bayes factor:
(hypotheses) as a whole —the best model in the set may                                  / BXΛ << 1.
merely be the best of in a set of poor models. Knowing
                                                                       To genuinely doubt the baseline model, ΛCDM, it is
that the best model in the current set of models is not
                                                                   not sufficient that R > 1, but additionally, the probability
particularly good model would provide motivation to
                                                                   of ΛCDM must also decrease such that its posterior
search for a better model, and hence may lead to model
                                                                   probability is greater than its prior probability, i.e.
discovery.
                                                                   p(Λ|d) < p(Λ). We can define:
    One way of assigning some measure of the absolute
                                                                                       RΛ ≡ p(Λ|d) / p(Λ).
goodness of a model is to use the concept of Bayesian
doubt, first introduced by [67]. Bayesian doubt works                  For ΛCDM to be doubted, the following two
by comparing all the known models in a set with an                 conditions must be fulfilled:
idealized model, which acts as a benchmark model.                                        R > 1, RΛ < 1.
    An application of the Bayesian doubt method for the                If these two conditions are fulfilled, then it suggests
cosmological model building is given in [44, 45]. One              that the set of known models is incomplete, and gives
of the most important questions in cosmology is to                 motivation to search for a better model not yet included,
identify the fundamental model underpinning the vast               which may lead to model discovery.



                                                             113
    In [67] a way of computing an absolute upper bound               In [54] the following elements related to hypothesis
for p(d|X) achievable among the class of known models            driven science are conceptualized: a phenomenon
has been proposed. Finally it was found that current             observed, a model interpreting this phenomenon, the
cosmic microwave background (CMB), matter power                  metadata defining the related computation together with
spectrum (mpk) and Type Ia supernovae (SNIa)                     the simulation definition (for simulation a declarative
observations.do not require the introduction of an               logic-based language is proposed). Specific attention in
alternative model to the baseline ΛCDM model. The                this work is devoted to hypothesis definition. The
upper bound of the Bayesian evidence for a presently             explanation a scientific hypothesis conveys is a
unknown dark energy model against ΛCDM gives only                relationship between the causal phenomena and the
weak evidence in favor of the unknown model. Since               simulated one, namely, that the simulated phenomenon
this is an absolute upper bound, it was concluded that           is caused by or produced under the conditions set by the
ΛCDM remains a sufficient phenomenological                       causal phenomena. By running the simulations defined
description of currently available observations.                 by the antecedents in the causal relationship, the
                                                                 scientist aims at providing hypothetical analysis of the
4 Facilities for the scientific hypothesis-                      studied phenomenon.
driven experiment support                                            Thus, the scientific hypothesis becomes an element
                                                                 of the scientific model that may replace a phenomenon.
4.1 Conceptualization of scientific experiments                  When computing a simulation based on a scientific
                                                                 hypothesis, i.e. according to the causal relationship it
    DIS      increasingly    becomes    dependent     on         establishes, the output results may be compared against
computational resources to aid complex researches. It            phenomenon observations to assess the quality of the
becomes paramount to offer scientists mechanisms to              hypothesis. Such interpretation provides for bridging the
manage the variety of knowledge produced during such             gap between qualitative description of the phenomenon
investigations. Specific conceptual modeling facilities          domain (scientific hypotheses may be used in
[54] are investigated to allow scientists to represent           qualitative (i.e., ontological) assertions) and the
scientific hypotheses, models and associated                     corresponding quantitative valuation obtained through
computational or simulation interpretations which can            simulations. According to the approach [54], complex
be compared against phenomena observations (Fig. 3).             scientific models can be expressed as the composition of
The model allows scientists to record the existing               computation models similarly to database views.
knowledge about an observable investigated
phenomenon, including a formal mathematical                      4.2 Hypothesis space browsers
interpretation of it, if any. Model evolution and model
sharing need also to be supported taking either a                    In the HyBrow (Hypothesis Space Browser) project
mathematical or computational view (e.g., expressed by           [58] the hypotheses for the biology domain are
scientific workflows). Declarative representation of             represented as a set of first-order predicate calculus
scientific model allows scientists to concentrate on the         sentences. In conjunction with an axiom set specified as
scientific issues to be investigated. Hypotheses can be          rules that model known biological facts over the same
used also to bridge the gap between an ontological               universe, and experimental data, the knowledge base
description of studied phenomena and the simulations.            may contradict or validate some of the sentences in
Conceptual views on scientific domain entities allow for         hypotheses, leaving the remaining ones as candidates
searching for definitions supporting scientific models           for new discovery. As more experimental data is
sharing among different scientific groups.                       obtained and rules identified, discoveries become
                                                                 positive facts or are contradicted. In the case of
    In [23] the engineering of hypothesis as linked data
                                                                 contradictions, the rules that caused the problems must
is addressed. A semantic view on scientific hypotheses
                                                                 be identified and eliminated from the theory formed by
shows their existence apart from a particular statement
                                                                 the hypotheses. In such model-theoretical approach, the
formulation in some mathematical framework. The
                                                                 validation of hypotheses considers the satisfiability of
mathematical equation is considered as not enough to
                                                                 the logical implications defined in the model with
identify the hypothesis, first because it must be
                                                                 respect to an interpretation. This might be applicable
physically interpreted, second because there can be
                                                                 also for simulation-based research, in which validation
many ways to formulate the same hypothesis. The link
                                                                 is decided based on the quantitative analysis between
to a mathematical expression, however, brings to the
                                                                 the simulation results and the observations [54].
hypothesis concept higher semantic precision. Another
                                                                 HyBrow is based on an OWL ontology and application-
link, in addition, to an explicit description of the
                                                                 level rules to contradict or validate hypothetical
explained phenomenon (emphasizing its “physical
                                                                 statements. HyBrow provides for designing hypotheses,
interpretation") can bring forth the intended meaning.
                                                                 and evaluating them for consistency with existing
By dealing with that hypothesis as a conceptual entity,
                                                                 knowledge, uses an ontology of hypotheses to represent
the scientists make it possible to change its statement
                                                                 hypotheses in machine understandable form as relations
formulation or even to assert a semantic mapping to
                                                                 between objects (agents) and processes [65].
another incarnation of the hypothesis in case someone
else reformulates it.                                                As an upgrade of HyBrow, the HyQue [12]
                                                                 framework adopts linked data technologies and employs
                                                                 Bio2RDF linked data to add to HyBrow semantic



                                                           114
interoperability       capabilities.    HyBrow/HyQue's            stored, or all the websites of tools relevant to
hypotheses are domain-specific statements that correlate          Neuroscience. Annotation may be structured or
biological processes (seen as events) in the First-Order          unstructured. Structured annotation means attaching a
Logic (FOL). Hypotheses are formulated as instances of            Concept (tag or term) to an Assertion. Unstructured
the HyQue Hypothesis Ontology and are evaluated                   annotation means attaching free text. Concepts are
through a set of SPARQL queries against biologically-             nodes in controlled vocabularies, which may also be
typed OWL and HyBrow data. The query results are                  hierarchical (taxonomies).
scored in terms of how the set of events correspond to
background expectations. A score indicates the level of           4.3 Scientific hypothesis formalization
support the data lends the hypothesis. Each event is                  An example showing on Fig. 6 the diversity of the
evaluated independently in order to quantify the degree           components of a scientific hypothesis model has been
of support it provides for the hypothesis posed.                  borrowed from the applications in Neuroscience
Hypothesis scores are linked as properties to the                 [54, 55] and in a human cardiovascular system in
respective hypothesis.                                            Computational       Hemodynamics         [23, 56].   The
    OBI (the Ontology for Biomedical Investigations)              formalization of a scientific hypothesis was provided by
project (http://obi-ontology.org) aims to model the               a mathematical model, by a set of differential equations
design of an investigation: the protocols, the                    for continuous processes, quantifying the variations of
instrumentation, and materials used in experiments and            physical quantities in continuous space-time and by the
the data generated [20]. Ontologies such as EXPO and              mathematical solver (HEMOLAB) for discrete
OBI enable the recording of the whole structure of                processes. The mathematical equations were represented
scientific investigations: how and why an investigation           in MathML, enabling models interchange and reuse.
was executed, what conclusions were made, the basis                   In [3] the formalism of quantitative process models
for these conclusions, etc. As a result of these generic          is presented that provides for encoding of scientific
ontology development efforts, the Minimum                         models formally as a set of equations and informally in
Information about a Genotyping Experiment (MIGen)                 terms of processes expressing those equations. The
recommends the use of terms defined in OBI. The use               model revision works as follows. For input it is required
of a generic or a compliant ontology to supply terms              an initial model; a set of constraints representing
will stimulate cross-disciplinary data-sharing and reuse.         acceptable changes to the initial model in terms of
As much detail about an investigation as possible in              processes; a set of generic processes that may be added
order to make the investigation more reproducible and             to the initial model; observations to which the revised
reusable can be collected [39].                                   model should fit. These data provide the approach with
    Hypothesis modeling is embedded into the                      a heuristic that guides search toward parts of the model
knowledge infrastructures being developed in various              space that are consistent with the observations. The
branches of science. One example of such infrastructure           algorithm generates a set of revised models that are
is considered under the name SWAN – a SemanticWeb                 sorted by their distance from the initial model and
Application in Neuromedicine [20]. SWAN is a project              presented with their mean squared error on the training
for developing an integrated knowledge infrastructure             data. The distance between a revised model and the
for the Alzheimer disease (AD) research community.                initial model is defined as the number of processes that
SWAN incorporates the full biomedical research                    are present in one but not in the other. The abilities of
knowledge lifecycle in its ontological model, including           the approach have been successfully checked in several
support for personal data organization, hypothesis                environmental domains.
generation,      experimentation,     laboratory     data             Formalisms for hypothesis formation are mostly
organization, and digital pre-publication collaboration.          monotonic and are considered to be not quite suitable
The common ontology is specified in an RDF Schema.                for knowledge representation, especially in dealing with
SWAN’s content is intended to cover all stages of the             incomplete knowledge, which is often the case with
“truth discovery” process in biomedical research, from            respect to biochemical networks. In [69] knowledge
formulation of questions and hypotheses, to capture of            based framework for the general problem of hypothesis
experimental data, sharing data with colleagues, and              formation is presented. The framework has been
ultimately the full discovery and publication process.            implemented by extending BioSigNet-RR – a
    Several information categories created and managed            knowledge based system that supports elaboration
in SWAN are defined as subclasses of Assertion. They              tolerant representation and non-monotonic reasoning.
include Publication, Hypothesis, Claim, Concept,                  The main features of the extended system provide: (1)
Manuscript, DataSet, and Annotation. An Assertion                 seamless integration of hypothesis formation with
may be made upon any other Assertion, or upon any                 knowledge representation and reasoning; (2) use of
object specifiable by URL. For example, a scientist can           various resources of biological data as well as human
make a Comment upon, or classify, the Hypothesis of               expertise to intelligently generate hypotheses; (3)
another scientist. Linking to objects “outside” SWAN              support for ranking hypotheses and for designing
by URL allows one to use SWAN as metadata to                      experiments to verify hypotheses. The extended system
organize – for example – all one’s PDFs of publications,          is positioned as a prototype of an intelligent research
or the Excel files in which one’s laboratory data is              assistant of molecular biologists.




                                                            115
                                  Fig. 6. Elements of the scientific hypothesis model

                                                                   study functional genomics in the yeast Saccharomyces
4.4 Hypothesis-driven robots
                                                                   cerevisiae, specifically to identify the genes encoding
    The Robot Scientist [66] oriented on genomic                   'locally orphan enzymes'. Adam uses a comprehensive
applications is a physically implemented system which              logical model of yeast metabolism, coupled with a
is capable of running cycles of scientific                         bioinformatic database (Kyoto Encyclopaedia of Genes
experimentation and discovery in a fully automatic                 and Genomes – KEGG) and standard bioinformatics
manner: hypothesis formation, experiment selection to              homology search techniques (PSI-BLAST and FASTA)
test these hypotheses, experiment execution using                  to hypothesize likely candidate genes that may encode
robotic system, results analysis and interpretation,               the locally orphan enzymes. This hypothesis generation
repeating the cycle (closed-loop in which the results              process is abductive.
obtained are used for learning from them and feeding                    To formalize Adam’s functional genomics
the resulting knowledge back into the experimental                 experiments, the LABORS ontology (LABoratory
models). Deduction, induction and abduction are types              Ontology for Robot Scientists) has been developed.
of logical reasoning used in scientific discovery (section         LABORS is a version of the ontology EXPO (as an
3). The full automation of science requires 'closed-loop           upper layer ontology) customized for Robot scientists to
learning', where the computer not only analyses the                describe biological knowledge. LABORS is expressed
results, but learns from them and feeds the resulting              in OWL-DL. LABORS defines various structural
knowledge back into the next cycle of the process                  research units, e.g. trial, study, cycle of study and
(Fig. 6).                                                          replicate as well as design strategy, plate layout,
    In the Robot Scientist the automated formation of              expected actual results. The respective concepts and
hypotheses is based on the following key components:               relations in the functional genomics data and metadata
    1. Machine–computable representation of the                    are also defined. Both LABORS and the corresponding
domain knowledge.                                                  database (used for storing the instances of the classes)
                                                                   are translated into Datalog in order to use the
    2. Abductive or inductive inference of novel                   SWI-Prolog reasoner for required applications [39].
hypotheses.
                                                                        There were two types of hypotheses generated. The
    3. An algorithm for the selection of hypotheses.               first level links an orphan enzyme, represented by its
    4. Deduction of the experimental consequences of               enzyme class (E.C.) number, to a gene (ORF) that
hypotheses.                                                        potentially encodes it. This relation is expressed as a two
    Adam, the first Robot Scientist prototype, was                 place predicate where the first argument is the ORF and the
designed to carry out microbial growth experiments to



                                                             116
second the E.C. number. An example of hypothesis at this                                      based on graph theory and logical modeling makes it
level is: encodesORFtoEC('YBR166C', '1.1.1.25').                                              possible to keep an accurate track of all the result units
     The second level of hypothesis involves the association                                  used for different goals, while preserving the semantics
between a specific strain, referenced via the name of its                                     of all the experimental entities involved in all the
missing ORF, and a chemical compound which should                                             investigations. It is shown how experimentation and
affect the growth of the strain, if added as a nutrient to its                                machine learning are used to identify additional
environment. This level of hypothesis is derived from the                                     knowledge to improve the metabolic model [13].
first by logical inference using a specific model of yeast
                                                                                              4.5 Hypotheses as data in probabilistic databases
metabolism. An example of such a hypothesis is: affects
growth('C00108','YBR166C'), where the first argument is                                           Another view of hypotheses encoding and
the compound (names according to KEGG) and the second                                         management is presented in [51]. Authors use
argument is the strain considered.                                                            probabilistic database techniques for hypotheses
     Adam then designs the experimental assays required                                       systematic construction and management. MayBMS
to test these hypotheses for execution on the laboratory                                      [30], a probabilistic database management system, is
robotic system. These experiments are based on a two-                                         used as a core for hypothesis management. This
factor design that compares multiple replicates of the                                        methodology (called γ-DB) enables researchers to
strains with and without metabolites compared against                                         maintain several hypotheses explaining some
wild type strain controls with and without metabolites.                                       phenomena and provides evaluation mechanism based
                                                                                              on Bayesian approach to rank them.
System model and        Initial START point:                    Experiment generation
 knowledge base        Hypotheses generation                         and design                   The construction of γ-DB database comprises
                                                                                              several steps. In the first step, phenomenon and
                                                                    Execution of              hypothesis entities are provided as input to the system.
                                               Cycles of          experiments on an
         Update the
                                         automated hypotheses     automated robotic
                                                                                              Hypothesis is a set of mathematical equations expressed
        system model                         generation and
                                            experimentation
                                                                       system                 as functions in W3C MathML-based format and is
                                                                                              associated with one or more simulation trial dataset,
                                                                                              consisting of tuples with input variables of equation and
                        Analysis of results                Collection of experimental
New knowledge            by statistics and                  observations and other            its corresponding output as functionally dependent (FD)
                        machine learning                           meta-data
                                                                                              variables (the predictions). Phenomenon is represented
          Fig. 6. Hypothesis driven closed-loop learning                                      by at least one empirical dataset similar to simulation
                                                                                              trials. In the next step, the system deals with hypotheses
    Adam follows a hypothetico-deductive methodology
                                                                                              and phenomena in the following way: 1) researcher has
(section 2). Adam abductively hypothesizes new facts
                                                                                              to provide some meta data about hypotheses and
about yeast functional biology, then it deduces the
                                                                                              phenomena; e.g., hypotheses need to be associated with
experimental consequences of these facts using its
                                                                                              the respective phenomena and assigned a prior
model of metabolism, which it then experimentally
                                                                                              confidence distribution (uniform by default according to
tests. To select experiments Adam takes into account
                                                                                              the principle of maximum entropy (3.2.3)); 2) functional
the variable cost of experiments, and the different
                                                                                              dependencies (FD) are extracted from equations in order
probabilities of hypotheses. Adam chooses its
                                                                                              to obtain database schema to store simulations and
experiments to minimize the expected cost of
                                                                                              experimental data; it should be mentioned that to
eliminating all but one hypothesis. This is in general a
                                                                                              precisely identify hypothesis formulation the special
NP complete problem and Adam uses heuristics to find
                                                                                              attributes for phenomena and hypothesis references are
a solution [65].
                                                                                              introduced into FD; 3) tuples are synthesized from
    It is now likely that the majority of hypotheses in                                       simulation trials and observational data by uncertain
biology are computer generated. Computers are                                                 pseudo-transitive closure and reasoning; 4) finally, the
increasingly automating the process of hypothesis                                             probabilistic γ-DB database is formed. Once
formation, for example: machine learning programs                                             phenomenon and hypothesis (with empirical datasets
(based on induction) are used in chemistry to help                                            and simulation trials) are produced it becomes possible
design drugs; and in biology, genome annotation is                                            to manipulate them with database tools.
essentially a vast process of (abductive) hypothesis
                                                                                                  MayBMS provides tools to evaluate competing
formation. Such computer-generated hypotheses have
                                                                                              hypotheses for the explanation of a single phenomenon.
been necessarily expressed in a computationally
                                                                                              With prior probabilities already provided the system
amenable way, but it is still not common practice to
                                                                                              allows to make one or more (if new observational data
deposit them into a public database and make them
                                                                                              appears) Bayesian inference steps. In each step the prior
available for processing by other applications [65].
                                                                                              probability is updated to posterior according to Bayes’
    The details describing the software and informatics                                       theorem. As a result, hypotheses which better explain
decisions in the Robot Scientist project can be found in                                      phenomenon get higher probabilities enabling
[65, 66] and online at the website http://www.aber.ac.uk/                                     researchers to make more confident decisions (see also
compsci/Research/bio/robotsci/data/informatics/). The                                         3.2.1). The γ-DB approach provides a promising way to
details for developing the formalization used for                                             analyse hypotheses in large scale DIR as uncertain
Adam’s functional genomics investigations can be                                              predictive database in face of empirical data.
found in [13, 39]. An ontology-based formalization




                                                                                        117
5 Examples of hypothesis-driven scientific                         mapping of all            neural    connections within
                                                                   an organism's nervous system. The production and study
research                                                           of connectomes, known as connectomics, may range in
                                                                   scale from a detailed map of the full set of neurons and
5.1 Besançon Galaxy model
                                                                   synapses within part or all of the nervous system of an
    Various models in astronomy heavily rely on                    organism to a macro scale description [15] of the
hypotheses. One of the most impressive is the Besançon             functional and structural connectivity between all
galaxy model (BGM) [16, 60, 61] evolving for many                  cortical areas and subcortical structures. The ultimate
years and representing the population and structure                goal of connectomics is to map the human brain. In
synthesis model for the Milky Way. It allows                       functional magnetic resonance imaging (fMRI),
astronomers to test hypotheses on the star formation               associations are thought to represent functional
history, star evolution, and chemical and dynamical                connectivity, in the sense that the two regions of the
evolution of the Galaxy. From the beginning, the aim of            brain participate together in the achievement of some
the BGM was not only to be able to simulate reasonable             higher-order function, often in the context of performing
star counts but further to test scenarios of Galactic              some task. fMRI has emerged as a powerful tool used to
evolution from assumptions on the rate of star formation           interrogate a multitude of functional circuits
(SFR), initial mass function (IMF), and stellar                    simultaneously. This has elicited the interest of
evolution.                                                         statisticians working in that area. At the level of basic
    We will further focus on the renewed BGM [16], in              measurements, neuroimaging data can be considered to
which authors draw their attention to the Galaxy thin              consist typically of a set of signals (usually time series)
disk treatment and use of Tycho-2 as a testing dataset.            at each of a collection of pixels (in two dimensions) or
The parameters of BGM (such as IMF, SFR and                        voxels (in three dimensions). Building from such data,
evolutionary track sets) explicitly and model ingredients          various forms of higher-level data representations are
implicitly can be treated as hypotheses. Model                     employed in neuroimaging. In recent years a substantial
ingredients include the treatment of binarity, the local           interest in network-based representations has emerged
stellar mass densities of thin disk, extinction model,             in neuroimaging to use networks to summarize
age-metallicity and age-velocity relations, radial scale           relational information in a set of measurements,
length, the age of the Galaxy thin disc, different sets of         typically assumed to be reflective of either functional or
the star atmosphere models, etc.                                   structural relationships between regions of interest in
                                                                   the brain. With neuroimaging now a standard tool in
    Tycho-2 dataset and χ2-type statistics test is used to         clinical neuroscience, quickly moving towards a time in
test various versions of these hypotheses in order to              which we will have available databases composed of
choose the most appropriate ones and update model to               large collections of secondary data in the form of
better fit the provided data. The tests were made by               network-based data objects is predictable.
comparing star counts and (B−V)T colour distributions
between data and simulations. Two different tests were                 One of the most basic tasks of interest in the analysis
used to evaluate the adequacy of the stellar densities             of such data is the testing of hypotheses, in answer to
globally and to test the shape of the colour distribution.         questions such as “Is there a difference between the
                                                                   networks of these two groups of subjects?" Networks
    Due to the fact, that some ingredients of the model            are not Euclidean objects, and hence classical methods
are highly correlated (such as the IMF, SFR and the                of statistics do not directly apply. Network-based
local mass density) the authors defined default models             analogues of classical tools for statistical estimation and
as a combination of a new set of ingredients that                  hypothesis testing are investigated [21, 22]. Such
significantly improve the fit to Tycho data. So, 11 IMF            research is motivated by the 1000 Functional
functions, 2 SFR functions, 2 evolutionary track sets, 3           Connectomes Project (FCP) launched in 2010 [7].The
sets of atmosphere models, 3 values for the age of the             1000 FCP [74] constitutes the largest data set of its kind
formation of the thin disk, 3 sets of values of the thin           similarly to large data sets in genetics. Other projects
disk local stellar volume mass density were tested. As a           (such as the Human Connectome Project (HCP)) are
result of testing, the two most appropriate IMS and SFR            aimed to build a network map of the human brain in
hypotheses were chosen. Based on this experience, an               healthy, living adults. The total volume of data
investigation of the thick disc is underway using SDSS             produced by the HCP will likely be multiple
and 2MASS surveys.                                                 petabytes [46]. HCP informatics platform includes data
                                                                   management system ConnectomeDB that is based on
5.2 Connectome analysis based on network data
                                                                   the XNAT imaging informatics platform [47], a widely
    In the neuroscience community the development of               used open source system for managing and sharing
common paradigms for interrogating the myriad                      imaging and related data.
functional systems in the brain remains to be the core                 Visualization, processing and analysis of high-
challenge. Building on the term “connectome,” coined               dimensional data such as images often requires some
to describe the comprehensive map of neural                        kind of preprocessing to reduce the dimensionality of
connections in the human brain, the “functional                    the data and find a mapping from the original
connectome” denotes the collective set of functional               representation to a low-dimensional vector space. The
connections in the human brain (its “wiring diagram”)              assumption is that the original data resides in a low-
[7]. More broadly, a connectome would include the



                                                             118
dimensional subspace or manifold [11], embedded in                     were significant to influence patterns of brain
the original space. This topic of research is called                   connectivity. The null hypothesis of no group
dimensionality reduction, non-linear dimensionality                    differences was rejected with high probability. Similarly
reduction, including methods for parameterization of                   for the three different age cohorts the null hypothesis of
data using low-dimensional manifolds as models.                        no cohort differences also was rejected with high
Within the neural information processing community                     probability.
this has become known as manifold learning. Methods                        On such examples it was shown [21] that the
for manifold learning are able to find non-linear                      proposed global test has sufficient power to reject the
manifold parameterizations of datapoints residing in                   null hypothesis in cases when mass-univariate approach
high-dimensional spaces, very much like Principal                      (considered to be the gold standard in fMRI research
Component Analysis (PCA) is able to learn or identify                  [43]) fails to detect the differences at the local level.
the most important linear subspace of a set of data                    According to the mass-univariate approach statistical
points (projecting data on a n-dimensional linear                      analysis is performed iteratively on all voxels to identify
subspace which maximizes the variance of the data in                   brain regions whose fMRI detected responses display
the new space).                                                        significant statistical effects. Thus it was shown that a
    In [21] necessary mathematical properties associated               framework for network-based statistical testing is more
with a certain notion of a ‘space’ of networks used to                 statistically powerful, than a mass-univariate approach.
interpret functional neuroimaging connectome-oriented                      It is expected that in the near future there will be a
data are established. Extension of the classical statistics            plethora of databases of network-based objects in
tools to network-based datasets, however, appeared to                  neuroscience motivating the development and extension
be highly non-trivial. The main challenge in such an                   of various tools from classical statistics to global
extension is due to the fact that networks are not                     network data.
Euclidean objects (for which classical methods were
                                                                           In the [70] paper discussion the relationship between
developed) – rather, they are combinatorial objects,
                                                                       neuroimaging and Big Data areas it is analyzed how
defined through their sets of vertices and edges. In [21]
                                                                       modern       neuroimaging       research   represents     a
it was shown that networks can be associated with
                                                                       multifactorial and broad ranging data challenge,
certain natural subsets of Euclidean space, and
                                                                       involving the growing size of the data being acquired;
demonstrated that through a combination of tools from
                                                                       sociological and logistical sharing issues; infrastructural
geometry, probability on manifolds, and high-
                                                                       challenges for multi-site, multi-datatype archiving; and
dimensional statistical analysis it is possible to develop
                                                                       the means by which to explore and mine these data. As
a principled and practical framework in analogy to
                                                                       neuroimaging advances further, e.g. aging, genetics, and
classical tools. In particular, an asymptotic framework
                                                                       age-related disease, new vision is needed to manage and
for one- and two-sample hypothesis testing has been
                                                                       process this information while marshalling of these
developed. Key to this approach is the correspondence
                                                                       resources into novel results. It is predicted that on this
between an undirected graph and its Laplacian, where
                                                                       way “big data” can become “big” brain science.
the latter is defined as a matrix (associating with a
network). Graph Laplacian appeared to be particularly                  5.3 Climate in Australia
appropriate to be used for such matrices. The space of
graph Laplacians is used working in certain subsets of                     Another view on hypothesis representation and
Euclidian space which are some submanifolds of the                     evaluation is presented in [41]. Authors argue, that as
standard Euclidian space.                                              long as in DIS data relevant to some hypotheses gets
    The 1000 FCP describes functional neuroimaging                     continuously aggregated as time passes, hypotheses
data from 1093 subjects, located in 24 community-based                 should be represented as programs that are executed
centers. The mean age of the participants is 29 years, and             repeatedly, as new relevant amounts of data gets
all subjects were 18 years-old or older. It is of interest to          aggregated. Their method and techniques are illustrated
compare the subject-specific networks of males and                     by examining hypotheses about temperature trends in
females in the 1000 FCP data set. In [21] for the 1000                 Australia during the 20th century. The hypothesis being
FCP database comparing of networks with respect to the                 tested comes from [42], stated that the temperature
sex of the subjects, over different age group, and over                series is not stationary and is integrated of order 1 (I(1)).
various collection sites is considered. It is shown that it is         Non-stationarity means that the level of the time series
necessary to compute the means in each subgroup of                     is not stable in time and can show increasing and
networks. This was done by constructing the Euclidean                  decreasing trends. I(1) means that by differentiating the
mean of the Laplacians for each group of subjects in                   stochastic process a stationary process (main statistical
different age groups. Such group-specific mean                         properties of the series remain unchanged) is obtained.
Laplacians can then be interpreted as the mean functional              Phillips-Perron test and the Kwiatkowski–Phillips–
connectivity in each group. Such approach provides for                 Schmidt–Shin (KPSS) test are used and both of them
building the hypothesis tests about the average of                     are executed in R. Several data sources are crawled: 1)
networks or groups of networks to investigate the effect               The National Oceanographic and Atmospheric
of sex differences on entire networks.                                 Administration marine and weather information, 2)
                                                                       Australian Bureau of Meteorology dataset. The
    For the 1000 FCP data set it was tested using the                  framework consists of R interpreter and R SPARQL,
two-sample test for Laplacians whether sex differences                 tseries packages. Authors also used agINFRA for



                                                                 119
computation and rich semantics to support traditional              formulation applying logical reasoning, various methods
scientific workflows for natural sciences. Authors                 for hypothesis modeling and testing (including classical
received further evidence on different independent                 statistics, Bayesian hypothesis and parameter estimation
dataset that time series is integrated of order 1.                 methods, hypothetico-deductive approaches)             are
                                                                   briefly introduced. Special attention is given to
5.4 Financial market                                               discussion of the data mining and machine learning
    Efficient-market hypothesis (EMH) is one of the                methods role in process of generation, selection and
most prominent in finance and “asserts that financial              evaluation of hypotheses as well as the methods for
markets are "informationally efficient"”. In [8] authors           motivation of new hypothesis formulation. Facilties of
test the weak form of EMH, stating that prices on traded           informatics for support of hypothesis-driven
assets (e.g., stocks, bonds, or property) already reflect          experiments, considered in the paper, are aimed at the
all past publicly available information. The null                  conceptualization of scientific experiments, hypothesis
hypothesis states that successive prices changes are               formulation and browsing in various domains (including
independent (random walk). The alternative hypothesis              biology, biomedical investigations, neuromedicine,
states that they are dependent. To check if the                    astronomy), automatic organization of hypothesis-
successive closing prices are dependent of each other              driven experiments. Examples of scientific researches
the following statistical tests were used: a serial                applying hypotheses considered in the paper include
correlation test, a runs test, an augmented Dickey-Fuller          modeling of population and structure synthesis of the
test and the multiple variance ratio test. Tests were              Galaxy, connectome-related hypothesis testing,
performed on daily closing prices from the six European            studying of temperature trends in Australia, analysis of
stock markets (France, Germany and UK, Greece,                     stock markets applying the EMN (Efficient market
Portugal and Spain) during the period between 1993 and             hypothesis), as well as algorithmic generation of
2007. The result of each test states whether successive            hypotheses in the IBM Watson project applying the
closing prices are dependent of each other.                        NLP and knowledge representation and reasoning
                                                                   technologies. An introduction into the state of the art of
    Test provides evidence that for monthly prices and             the hypothesis-driven research presented in the paper
returns the null hypothesis should not be rejected for all         opens a way for investigation of the generalized
six markets. If daily prices are concerned the null                approaches for efficient organization of hypothesis-
hypothesis is not rejected for France, Germany, UK and             driven experiments applicable for various branches
Spain, but this hypothesis is rejected for Greece and              of DIS.
Portugal. However, on the 2003-2007 dataset the null
hypothesis for these two countries is not rejected as
well.                                                              References
    In [8] Bollen et al. use different approach to test                [1] Agresti, A., Finlay, B. Statistical Methods for
EMH. Authors investigate whether public sentiment, as                      the Social Sciences (4th Edition), 2008. –
expressed in large-scale collections of daily Twitter                      P. 624.
posts, can be used to predict the stock market. They                   [2] Alferes, J. J., Pereira, L. M., Swift, T.
build public mood time series by sentiment analysis of                     Abduction in well-founded semantics and
tweets from February 28, 2008 to December 19, 2008                         generalized stable models via tabled dual
and try to show that it can predict Dow Jones Index                        programs. In: TPLP, 2004. – Vol. 4, No. 4. –
corresponding values. The null hypothesis states that the                  P. 383–428.
mood time series do not predict DJIA values. Granger                   [3] Asgharbeygi, N., Langley, P., Bay, S., Arrigo,
causality analysis in which Dow Jones values and mood                      K. Inductive revision of quantitative process
time series are correlated is used to test the null                        models. In: Ecological modelling – Vol. 194,
hypothesis. Granger causality analysis is used to                          No. 1. – P. 70–79.
determine if one time series can predict another time-                 [4] Bacon, F. The new organon. In: R. M.
series. Its results reject the null hypothesis and claim                   Hutchins, (ed.), Great books of the western
that public opinion is predictive of changes in DJIA                       world. The works of Francis Bacon. Chicago,
closing values.                                                            Encyclopedia Britannica, Inc., 1952 – Vol. 30.
                                                                           – P. 107–195.
6 Conclusion
                                                                       [5] Barber, D. Bayesian Reasoning and Machine
    The objective of this study is to analyze, collect and                 Learning. Cambridge University Press, 2010. –
systematize information on the role of hypotheses in the                   P. 720.
data intensive research process as well as on support of               [6] Bartha, P. Analogy and Analogical Reasoning.
hypothesis formation, evaluation, selection and                            In: The Stanford Encyclopedia of Philosophy,
refinement in course of the natural phenomena                              2013. – http://plato.stanford.edu/archives/
modeling and scientific experiments. The discussion is                     fall2013/entries/reasoning-analogy/
started with the basic concepts defining the role of                   [7] Biswal, B.B., Mennes, M., Zuo, X.N.,
hypotheses in the formation of scientific knowledge and                    Gohel, S., Kelly, C., Smith, S.M.,
organization of the scientific experiments. Based on                       Windischberger, C. Toward discovery science
such concepts, the basic approaches for hypothesis                         of human brain function. In: Proceedings of the




                                                             120
     National Academy of Sciences, 2010. – V. 107,              [21] Ginestet, C.E., Balanchandran, P.,
     No. 10. – P. 4734–4739.                                         Rosenberg, S., Kolaczyk, E.D. Hypothesis
 [8] Bollen, J., Mao, H., Zeng, X. Twitter mood                      Testing For Network Data in Functional
     predicts the stock market. In: Journal of                       Neuroimaging. In: arXiv preprint
     Computational Science, 2011. – V. 2, No. 1. –                   arXiv:1407.5525, 2014.
     P. 1–8.                                                    [22] Ginestet, C. E., Fournel, A. P., Simmons, A.
 [9] Borges, M. R. Efficient market hypothesis in                    Statistical network analysis for functional MRI:
     European stock markets. In: The European                        summary networks and group comparisons. In:
     Journal of Finance, 2010. – V. 16, No. 7. –                     Frontiers in computational neuroscience, 2014.
     P. 711–726.                                                     – Vol. 8.
[10] Breiman, L. Statistical Modeling: The Two                  [23] Gonçalves, B., Porto, F. A Lattice-Theoretic
     Cultures. In: Statistical Science, 2001. – V. 16,               Approach for Representing and Managing
     No. 3. – P. 199–231.                                            Hypothesis-driven Research. In: AMW, 2013.
[11] Brun, A. Manifold learning and representations             [24] Gonçalves, B., Porto, F., Moura, A. M. C. On
     for image analysis and visualization.                           the semantic engineering of scientific
     Department of Biomedical Engineering,                           hypotheses as linked data. In: Proceedings of
     Linköpings universitet, 2006.                                   the 2nd International Workshop on Linked
[12] Callahan, A., Duumontier, M., Shah, N.                          Science, 2012.
     HyQue: Evaluating hypotheses using Semantic                [25] Haber, J. Research Questions, Hypotheses, and
     Web technologies. In: J. Biomedical Semantics,                  Clinical Questions. In: Evolve Resources for
     2011. – V. 2, No. S-2. – P. S3.                                 Nursing Research, 2010. – P. 27–55.
[13] Castrillo, J.I. , S.G. Oliver (eds.). Yeast                [26] Hastie, T., Tibshirani, R., Friedman, J.,
     Systems Biology: Methods and Protocols. In:                     Franklin, J. The elements of statistical learning:
     Methods in Molecular Biology, Springer, 2011.                   data mining, inference and prediction. In: The
     – V. 759. – P. 535.                                             Mathematical Intelligencer, 2005. – Vol. 27,
[14] Citrigno, S., Eiter, T., Faber, W., Gottlob, G.,                No. 2. – P. 83–85.
     Koch, C., Leone, N., Scarcello, F. The dlv                 [27] Hawthorne, J. Inductive Logic. In: The
     system: Model generator and application                         Stanford Encyclopedia of Philosophy, 2014 –
     frontends. In: Proceedings of the 12th                          http://plato.stanford.edu/archives/sum2014/entri
     Workshop on Logic Programming, 1997. – P.                       es/logic-inductive/
     128–137.                                                   [28] Hempel, C. G. Fundamentals of concept
[15] Craddock, R.C., Jbabdi, S., Yan, C.G.,                          formation in empirical science. In: Int.
     Vogelstein, J.T., Castellanos, F.X., Di Martino,                Encyclopedia Unified Science, 1952. – V. 2,
     A., Milham, M.P. Imaging human connectomes                      No. 7.
     at the macroscale. In: Nature methods, 2013. –             [29] Hey, T., Tansley, S., Tolle, K. (eds.). The
     V. 10, No. 6. – P. 524–539.                                     fourth paradigm: Data-intensive scientific
[16] Czekaj, M.A., Robin, A.C., Figueras, F.,                        discovery. Redmond, Microsoft Research,
     Luri, X., Haywood, M. The Besançon Galaxy                       2009. – P. 252.
     model renewed I. Constraints on the local star             [30] Huang, J., Antova, L., Koch, C., Olteanu, D.
     formation history from Tycho data. In: arXiv                    MayBMS: a probabilistic database management
     preprint arXiv:1402.3257, 2014.                                 system. In: Proceedings of the 2009 ACM
[17] Dredze, M, Crammer, K., Pereira, F.                             SIGMOD International Conference on
     Confidence-Weighted Linear Classification. In:                  Management of data, 2009. – P. 1071–1074.
     Proceedings of the 25th International                      [31] IBM SPSS Statistics for Windows, Version
     Conference on Machine Learning, Helsinki,                       22.0. Armonk, NY: IBM Corp. IBM SPSS
     Finland, 2008. – P. 264-271.                                    Statistics base. IBM Corp., 2013.
[18] Ferrucci, D., Brown, E., Chu-Carroll, J., Fan, J.,         [32] Ihaka, R., Gentleman, R. R: a language for data
     Gondek, D., Kalyanpur, A.A., Welty, C.                          analysis and graphics. In: Journal of
     Building Watson: An overview of the DeepQA                      computational and graphical statistics, 1996. –
     project. In: AI magazine, 2010. –V. 31, No. 3. –                Vol. 5, No. 3. – P. 299–314.
     P. 59–79.                                                  [33] Inoue K., Sato T., Ishihata M., Kameya Y.,
[19] Field, A. Discovering statistics using IBM                      Nabeshima H. Evaluating abductive hypotheses
     SPSS statistics. In: Sage, 2013. – P. 915.                      using and EM algorithm on BDDs. In:
[20] Gao, Y., Kinoshita, J., Wu, E., Miller, E.,                     Proceedings of IJCAI-09, 2009. – P. 810–815.
     Lee, R., Seaborne, A., Clark, T. SWAN:                     [34] Ivezić, Ž., Connolly, A. J., VanderPlas, J. T.,
     A distributed knowledge infrastructure for                      Gray, A. Statistics, Data Mining, and Machine
     Alzheimer disease research. In: Web                             Learning in Astronomy: A Practical Python
     Semantics: Science, Services and Agents on the                  Guide for the Analysis of Survey Data.
     World Wide Web, 2006. - V. 4, No. 3. – P.                       Princeton University Press, 2014. – P. 552.
     222–228.



                                                          121
[35] Kakas, A.C., Michael, A., Mourlas, C. ACLP:              [48] McComas, W.F. The principal elements of the
     Abductive constraint logic programming. In:                   nature of science: dispelling the myths. In: The
     The Journal of Logic Programming, 2000. –                     Nature of Science in Science Education, 1998.
     Vol. 44, No. 1. – P. 129–177.                                 – P. 53–70.
[36] Kakas, A.C., Kowalski, R.A., Toni, F.                    [49] Menzies, T. Applications of Abduction:
     Abductive Logic Programming. In: Journal of                   Knowledge-Level Modeling, In: International
     Logic and Computation, 1993. – Vol. 2, No. 6.                 Journal of Human-Computer Studies, 1996. –
     – P. 719–770.                                                 V. 45, No. 3. – P. 305–335.
[37] Kerlinger, F.N., Lee, H.B. Foundations of                [50] Nickles, T. (ed.). Scientific discovery: Case
     behavioral research: Educational and                          studies. Taylor & Francis, 1980. – Vol. 2. –
     psychological inquiry. New York: Holt,                        P. 501.
     Rinehart and Winston, 1964. – P. 739.                    [51] Plotkin, G.D. A note on inductive
[38] King, R.D., Liakata, M., Lu, C., Oliver, S.G.,                generalization. In: Machine Intelligence.
     Soldatova, L.N. On the formalization and reuse                Edinburgh University Press, 1970. – Vol. 5. –
     of scientific research. In: Journal of The Royal              P. 153–163.
     Society Interface, 2011. – Vol. 8, No. 63. –             [52] Poincaré, Henri. The Foundations of Science:
     P. 1440–1448.                                                 Science and Hypothesis, The Value of Science,
[39] King, R.D., Whelan, K.E., Jones, F.M.,                        Science and Method. The Project Gutenberg
     Reiser, P.G., Bryant, C.H., Muggleton, S.H.,                  EBook, 2012. – Vol. 39713. – P 554.
     Oliver, S.G. Functional genomic hypothesis               [53] Popper, K.. The Logic of Scientific Discovery
     generation and experimentation by a robot                     (Taylor & Francis e-Library ed.). London and
     scientist. Nature, 2004. – Vol. 427, No. 6971. –              New York: Routledge / Taylor & Francis e-
     P. 247–252.                                                   Library, 2005.
[40] Lakshmana Rao, J R. Scientific 'Laws',                   [54] Porto, F. Big Data in Astronomy. The LIneA-
     'Hypotheses' and 'Theories'. In: Meanings and                 DEXL case. Presentation at the EMC Summer
     Distinctions. Resonance, 1998. – Vol. 3. – P.                 School on BIG DATA – NCE/UFRJ, 2013.
     69–74.                                                   [55] Porto, F. Moura, A. M. C., Gonçalves, B.,
[41] Lappalainen, J., Sicilia, M.Á., Hernández, B.                 Costa, R., Spaccapietra, S. A Scientific
     Automatic Hypothesis Checking Using                           Hypothesis Conceptual Model. In: ER
     eScience Research Infrastructures, Ontologies,                Workshops, 2012. – Vol. 7518. – P. 101–110.
     and Linked Data: A Case Study in Climate                 [56] Porto, F., Moura, A. M. C. Scientific
     Change Research. In: Procedia Computer                        Hypothesis Database. Report, 2011.
     Science, 2013. – Vol. 18. – P. 1172–1178.
                                                              [57] Porto, F., Spaccapietra, S. Data model for
[42] Lenten, L.J., Moosa, I.A. An empirical                        scientific models and hypotheses. In: The
     investigation into long-term climate change in                evolution of conceptual modeling, 2011. –
     Australia. In: Environmental Modelling &                      Vol. 6520. – P. 285–305.
     Software, 2003. – Vol. 18, No. 1. – P. 59–70.
                                                              [58] Racunas, S.A., Shah, N.H., Albert, I., Fedoroff,
[43] Mahmoudi, A., Takerkart, S., Regragui, F.,                    N.V. Hybrow: a prototype system for
     Boussaoud, D., Brovelli, A. Multivoxel Pattern                computer-aided hypothesis evaluation. In:
     Analysis for fMRI Data: A Review. In:                         Bioinformatics, 2004. – Vol. 20, No. 1. –
     Computational and mathematical methods in                     P. 257–264.
     medicine, 2012.
                                                              [59] Ray, O., Kakas, A. ProLogICA: a practical
[44] March, M.C. Advanced Statistical Methods for                  system for Abductive Logic Programming. In:
     Astrophysical Probes of Cosmology. In:                        Proceedings of the 11th International Workshop
     Springer Theses, 2013. – Vol. 20. – P. 177.                   on Non-monotonic Reasoning, 2006. – P. 304–
[45] March, M.C., Starkman, G.D., Trotta, R.,                      312.
     Vaudrevange, P. M. Should we doubt the                   [60] Robin, A.C., Reylé, C., Derrière, S., Picaud, S.
     cosmological constant?. In: Monthly Notices of                A synthetic view on structure and evolution of
     the Royal Astronomical Society, 2011. –                       the Milky Way. arXiv preprint astro-
     Vol. 410, No. 4. – P. 2488–2496.                              ph/0401052, 2004.
[46] Marcus, D.S., Harwell, J., Olsen, T., Hodge, M.,         [61] Robin, A., Crézé, M. Stellar populations in the
     Glasser, M.F., Prior, F., Van Essen, D.C.                     Milky Way-A synthetic model. In: Astronomy
     Informatics and data mining tools and strategies              and Astrophysics, 1986. – Vol. 157. – P. 71–90.
     for the human connectome project. In: Frontiers
                                                              [62] Rouder, J.N., Speckman, P.L., Sun, D.,
     in neuroinformatics, 2011. – Vol. 5.
                                                                   Morey, R.D., Iverson, G. Bayesian t tests for
[47] Marcus, D.S., Olsen, T.R., Ramaratnam, M.,                    accepting and rejecting the null hypothesis. In:
     Buckner, R.L. The extensible neuroimaging                     Psychonomic bulletin & review, 2009. –
     archive toolkit. In: Neuroinformatics, 2007. –                Vol. 16, No. 2. – P. 225–237.
     Vol. 5, No. 1. – P. 11–33.




                                                        122
[63] Schickore, J. Scientific Discovery. The                        hypothesis formation in biochemical networks.
     Stanford Encyclopedia of Philosophy, 2014 –                    In: Data Integration in the Life Sciences, 2005.
     http://plato.stanford.edu/archives/spr2014/entrie              – P. 121–136.
     s/scientific-discovery/                                   [70] Van Horn, J.D., Toga, A.W. Human
[64] Sivia, D.S., Skilling, J. Data Analysis. A                     neuroimaging as a “Big Data” science. In:
     Bayesian Tutorial. Oxford University Press                     Brain imaging and behavior, 2014. – Vol. 8,
     Inc., New York, 2006. – P. 264.                                No. 2. – P. 323–331.
[65] Soldatova, L.N., Rzhetsky, A., King, R. D.                [71] Van Nuffelen, B., Kakas, A. A-system:
     Representation of research hypotheses. In: J.                  Declarative programming with abduction. In:
     Biomedical Semantics, 2011. – Vol. 2, No. S-2.                 Logic Programming and Nonmotonic
     – P. S9.                                                       Reasoning, 2001. – P. 393–397.
[66] Sparkes, A., Aubrey, W., Byrne, E., Clare, A.,            [72] Weber, M. Experiment in Biology. The
     Khan, M. N., Liakata, M., King, R. D. Towards                  Stanford Encyclopedia of Philosophy, 2014. –
     Robot Scientists for autonomous scientific                     http://plato.stanford.edu/archives/fall2014/entri
     discovery. In: Autom Exp, 2010. – Vol. 2,                      es/biology-experiment/
     No 1.                                                     [73] Woodward, J. Scientific Explanation. The
[67] Starkman, G.D., Trotta, R., Vaudrevange, P.M.                  Stanford Encyclopedia of Philosophy, 2011. –
     Introducing doubt in Bayesian model                            http://plato.stanford.edu/archives/win2011/entri
     comparison. arXiv preprint arXiv:0811.2415,                    es/scientific-explanation/
     2008.                                                     [74] Yan, C.G., Craddock, R.C., Zuo, X.N., Zang,
[68] Tamaddoni-Nezhad, A., Chaleil, R., Kakas, A.,                  Y.F., Milham, M.P. Standardizing the intrinsic
     Muggleton, S.H. Application of abductive ILP                   brain: towards robust measurement of inter-
     to learning metabolic network inhibition from                  individual variation in 1000 functional
     temporal data. In: Machine Learning, 2006. –                   connectomes. In: Neuroimage, 2013. – Vol. 80.
     Vol. 64. – P. 209–230.                                         – P. 246–262.
[69] Tran, N., Baral, C., Nagaraj, V.J., Joshi, L.
     Knowledge-based integrative framework for




                                                         123