<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unveiling Latent States Behind Social Indicators</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emanuele Di Buccio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Lorenzet</string-name>
          <email>andrea.lorenzet@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Massimo Melucci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Neresini</string-name>
          <email>federico.neresini@unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Philosophy</institution>
          ,
          <addr-line>Sociology, Education and Applied Psychology</addr-line>
          ,
          <institution>University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The work reported in this paper aims at describing a project that leverages the potential of Information Retrieval and Machine Learning towards novel techniques that unveil the latent states of expert users such as sociologists and economists by means of indicators when the user is accessing large collection of newspapers, blogs, etc. An indicator measures the degree to which a certain latent state is present during interaction when exploring and searching an information repository. In this paper the state of a user is the particular condition that s/he is in at a speci c context with reference to a problematic issue induced by the data s/he accessed to; risk is an example of state and a risk indicator aims at providing a measure of the degree to which articles examined by the user evoke risk in her/his mind. Observable attributes, e.g. keywords, clickthrough data or links, are the input data to model states and compute indicators. In this work, starting from some results of a software architecture designed to support sociologists in investigating Techno-scienti c Issues in the Public Sphere, we will discuss some challenges and we will present a formal framework to address them where informative objects, e.g. news articles, states and attributes are uniformly modelled as vectors as it is customary in Information Retrieval or Machine Learning. This is our rst step towards a long term objective, i.e. generalizing the well known Learning to Rank framework towards a Learning to Search framework which would encompass multiple and simultaneous states.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Although Social Computing (SC) has a long history and dates back to, for
example, Sadowski's work [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the advent of big structured, semi-structured or
unstructured multimedia data motivates, from the one hand, computer scientists
to propose novel e cient and e ective methods to monitor and analyse social,
economic and political phenomena, thus developing a new research area called
Social Computing. On the other hand, politicians, sociologists and economists
are witnessing a major shift in social science research methodology thanks to
vast, complex arrays of data to work with [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Not only the amount of data is rapidly increasing, the variety of types of data
has also become larger and the degree of user interaction may be higher than
in the past. The typology of data that are accessed by social scientists, students
or causal users such as World Wide Web (WWW) surfers includes structured
data extracted from relational databases, semi-structured data received from
eXtended Markup Language (XML)-encoded streams such as news feed, or
unstructured data collected, indexed and delivered to end users by search engines
as inter-linked pages. The di erent data management systems mentioned above
can provide a variety of services to the end users who may search, browse and
annotate text, images, video and music.</p>
      <p>Together with users' behaviour data (e.g. click-through data or search
session), these systems can be utilised to organize and analyse the users' thoughts
and feelings about di erent controversial topics. For example, the vast and
complex arrays of data such as social media can thus be utilized to create
representations of techno-scienti c controversies which are able to trigger intense public
debates such as those concerning issues like climate change, Genetically Modi ed
Organisms (GMOs), nuclear power.</p>
      <p>
        When accessing to information by reading, browsing or annotating
multimedia documents, the end users form a view or judgement about something,
not necessarily based on fact or evidence, on the contrary, based on general
feeling or opinion. These situations not only clearly blur the traditional
boundaries among expertise, policy making, politics, and public opinion [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. They also
fuel an emotional state or reaction towards social issues and controversies. The
knowledge of the user's state is not only interesting in itself, it is also crucial
because it is such a state that often drives the user's decision about political
and economic behaviours in di erent contexts (e.g. purchases, elections or party
membership)[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        In this paper, the state of a user is therefore the particular condition that s/he
is in at a speci c context with reference to a problematic issue documented by the
data s/he accessed to. An example of user state is the condition that the user is
in with reference to themes related to the \risk society". The risk society refers to
the idea that within contemporary society Science and Technology (S&amp;T) issues
are imbued with fears and preoccupations about unforeseen e ects, calling for a
precautionary approach on the side of policy making, society, and the public [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
It follows that, a user who is reading a newspaper editorial may be in uenced
by the editor's opinion on a topical issue, s/he may perceive risk, con ict, worry
or controversy, and may collapse to the state in which s/he perceives some risk.
      </p>
      <p>Note that understanding the user's state is di erent from the understanding
the document's author opinion addressed by means of Sentiment and Opinion
Mining and Analysis systems which aim to classify the opinion of a document's
author who has implicitly expressed his opinion by means of a document. The
user's state is not necessarily encoded in words, on the contrary, it might be
understood by analysing di erent types of data ranging from click-through data
to natural language queries in combination with document attributes. Moreover,
our focus is not on users' mood that may be mined from short or tiny data such as
tweets or \likes". Our interest is on readers who might not immediately comment
and reveal their own feeling to the public, which is nevertheless of great interest
to expert users such as sociologists or economists who are asked to be acutely
aware of the issues of our world.</p>
      <p>This paper brie y reports on our advances in using SC techniques based on
Information Retrieval (IR) and Machine Learning (ML) to support sociologists in
investigating Techno-scienti c Issues in the Public Sphere (TIPS). Starting from
these initial results, it is our objective to further develop the framework and the
TIPS software architecture to unveil the latent states that a user may experience
when interacting with news articles that deal with techno-scienti c controversies.
In this context, IR may play a crucial role because it is naturally devoted to
meet user informative needs stemming from the problem of understanding social
phenomena, when these phenomena are encoded in large collections of
interlinked documents such as news, feeds, blogs, video and images.</p>
      <p>
        Besides extending TIPS, we have got a long term vision. Quoting Liu [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
search engines by now go \beyond the pure relevance-based ranking of documents
in their search results," since some prominent search engines \try to provide
rich presentation of search result to users. When the ranked list is no longer the
desired output, the learning-to-rank" { also known the application of ML to IR
{ \technologies need to be re ned: the change of the output space will naturally
lead to the change of the hypothesis space and the loss function, as well as
the change of the learning theory. On the other hand, the new search scenario
may be decomposed into several sub ranking tasks and many key components in
learning to rank can still be used. This may become a promising future work for
all the researchers currently working on learning to rank, which we would like
to call learning to search rather than learning to rank." In this paper, we name
this long term vision Learning To Search (LETS) where advanced systems are
designed to support complex search task where the query-response paradigm is
just a tactic within a search strategy. It is our opinion that considering multiple
and simultaneous user' states is a step forward LETS and as a consequence an
advance of SC technologies.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Motivations</title>
      <p>For some years, a project called TIPS3 has been carried out by us to develop,
experiment, and implement automatic procedures for collecting, classifying, and
analysing digital contents { mainly online news and user generated contents in
social media { in order to monitor S&amp;T issues and their evolution. TIPS collects
online articles from six newspapers and classi es them according to their
pertinence to S&amp;T topics; currently the corpus is constituted of more than a million
documents, collected since 2010. Fig. 1 reports an overview of the architecture
components.</p>
      <p>TIPS collects articles from online news such as RSS feeds associated to
speci c newspaper sections through a collector module. The articles are classi ed
and a risk indicator is calculated to measure the degree to which the articles
3 http://hal.cloud.tilaa.com/tips/
informative
objects sources
collector
module
classifiers and
indicators
indexing
module</p>
      <p>IR
technology</p>
      <p>user
interface
S&amp;T Classifier
Risk Indicator
Computation</p>
      <p>S&amp;T classifier score
S&amp;T classifier label</p>
      <p>Risk indicator
evoke risk in the users' mind. The articles are then enriched with the categories
associated by the classi ers and the indicators. Furthermore, the articles are
provided to an indexing module that stores the article attributes in a data repository
that provides e cient access though meta-data and content-based search using
IR technologies4. Additional layers can then be built on top of the IR module
to provide meta-data and content-based access through a WWW user interface
and to display dynamic charts representing indicator trends (Fig. 1).</p>
      <p>
        The remainder of this section brie y describes how the risk indicator is
currently instantiated in the current release of TIPS. A computer scientist may
nd the procedures implemented by this system rather rudimentary, however,
the manual intervention is limited and it was required by the sociologists who
wished to closely monitor the procedures as possible. A risk indicator relies on
a set of keywords manually identi ed by the sociologists through the support
of an unsupervised ML algorithm for the extraction of themes from an
unstructured corpus. The keywords were selected by (i) retrieving documents through
the query \risk", (ii) extracting topics and subtopics through an implementation
of the HPAM topic modelling algorithm [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and (iii) identifying terms in the
conceptual area of risk by the manual inspection of the sub-topics. For
example, a subset of the selected keywords could be: \infection", \danger", \alarm",
\catastrophic". Each document is thus represented in terms of these keywords,
as it is customary in many retrieval application domains.
4 The current version of TIPS relies on ElasticSearch: https://www.elastic.co
      </p>
      <p>An indicator described by the set of keywords K for the document set D, will
be computed as</p>
      <p>IK(D) =
1 X
jDj d2D</p>
      <p>
        IK(d)
IK(d) =
1 X nL(w; d)=B
jKj w2K nL(w; d)=B + K
(1)
(2)
where
and nL(w; d) is the frequency of the term w in the document d; nL(w; d) is
normalized by B: (1 b) + b avdgld(dl()C) where dl(d) is the length of the document d and
avgdl(C) is the average document length in the corpus C; b 2 [0; 1] is a
parameter that controls the weight assigned to the document length normalization. The
nL(w; d)=B normalization has been introduced in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The K's values control
the e ect of the term frequency on the indicator value for a document: the basic
idea is to consider a non-linear dependence between the frequency and the
indicator, thus delivering relatively high values already for small frequencies, and
then \saturates" for large term frequencies. The indicator obtained for a
document, IK(d), can be then used as additional document descriptor to perform
meta-data based search.
      </p>
      <p>In order to show a possible application of a risk indicator to actual news, in
the remainder of this section we will report on the trends obtained for the risk
indicator using an Support Vector Machine (SVM)-based classi er to identify
S&amp;T news articles. We considered a subset of articles collected by TIPS and
published from January 1st, 2010 to December 31, 2015; the total number of
articles constituting this subset is 630,549. A SVM-based classi er with linear
kernel was trained on a sample of the news corpus that was manually labelled by
a team of sociologists on the basis of their pertinence to S&amp;T issues. The
sample is constituted of 3817 documents; 1,393 were labelled as pertinent to S&amp;T,
while the remaining 2,424 as non pertinent. The e ectiveness of the classi er
was tested using a 60/40 split for training/test with 10 fold cross-validation. We
then classi ed all the 630,549 articles using the learned classi er. The risk
indicator was then computed for all the articles, for the subset of articles classi ed
as pertinent and for those classi ed as not-pertinent. The evolution of the risk
indicator on a monthly basis and in the time frame 2013-2015 is reported in
Fig. 2; since the time granularity is one month, the document set D in
Equation 1 for the lth time interval til, denoted by Dtil , is constituted by all the
articles published in the lth month of the time frame 2013-2015; in the event of
relevant and not-relevant trends, the document set is respectively constituted of
the relevant and the not relevant documents published in the time interval til.</p>
      <p>One of the limitations due to using only content based descriptors is evident
in Fig. 2: the trend of the risk indicator in the entire corpus di ers from that in
the subset of relevant articles, thus suggesting that the distribution of terms in
risky documents can di er from that in relevant documents; therefore, frequency
might be no longer the best evidence to capture the degree of risk.
Relevant</p>
      <p>Not Relevant</p>
      <p>All
0.06
0.05</p>
      <p>Fig. 3 reports the variation of the risk indicator when computed for
documents about \climate change" and documents about \nuclear power" for the
subset of the news articles collected in TIPS and published from 2013 to 2015;
also in this case the distribution of risky terms in the two subsets is di erent.
However, even if the current indicator is able to capture a di erence in terms
of perception of risk in \climate change" and \nuclear power" related articles,
the indicator is unable to explain why such di erence exists. A more suitable
representation of risk should be able to assists the users, e.g. specialists, also
in this task. The framework introduced in the remainder of this paper aims to
achieve such a goal through an explicit representation of states.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Challenges</title>
      <p>TIPS has been the starting point of the contribution described in this paper.
Since the early phases of the TIPS project, we realized that, beyond the
sociologists' requirements, there is a great potential and some challenges of SC {
mentioned in Section 1 { can be met if some current limitations of TIPS can be
overcome. The main limitations of TIPS that are considered in this paper are
the following ones:
Single state. Only one state (e.g. risk) can be at a time detected and measured
from the information objects provided as input. The source software in principle
is able to manage diverse states, but states are modelled as independent each
Climate Change</p>
      <p>Nuclear Power
0,06
0,05
0,04
R
O
T
A
ICD0,03
N
I
K
S
I
R
0,02
0,01
0
other. Moreover, some manual intervention is required to tailor the system to
manage a speci c state. The manual intervention to be made consists of
compiling controlled keyword vocabularies such as the vocabulary of keywords evoking
risk. Another state, e.g. \con ict" will require the compilation of another,
distinct keyword list, and a group of experts in con icts would be in charge of this
compilation.</p>
      <p>
        Content-based state modelling. The current implementation of TIPS only
exploits content-based descriptors such as keywords, dates, and class labels.
Because of the type of descriptors selected during system design, we have regarded
risk as similar to relevance5 and we used a normalized term frequency that
worked well in general content-based IR [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] to compute the risk indicator.
      </p>
      <p>However, in the light of detecting risk or other states, the current
computation of indicators soon appeared rather simplistic. Normalized term frequency is
not necessarily the best method to detect and measure risk and in general states
other than relevance. For example, the distribution of terms in risky documents
can di er from that in relevant documents and frequency might be no longer the
best evidence to capture the degree of risk and actual interaction with the end
user is ignored.
5 A document is relevant to an information need when it carries information that help
the user to meet his/her need</p>
      <p>
        In contrast, the users' perception of risk may be better captured by the user
behaviour during the interaction with the articles. Since an article might not
include \risk" although it could evoke risk, or may include the word while not
evoking the feeling, it is the complex of the users queries, click-through data,
and other interaction features that may suggest the \riskiness" of the articles.
Much information about the user's information need can be obtained through
interaction [
        <xref ref-type="bibr" rid="ref5 ref6">5,6</xref>
        ]. Interaction can be, for instance, used in the actual computation
of the indicator.
      </p>
      <p>No simulation-based indicators. TIPS has been mainly designed for
monitoring purposes by using IR and ML technologies. However, when studying the
relationship between techno-scienti c issues and the public's opinion, some tasks
could bene t from the possibility to simulate speci c scenarios, e.g. obtained by
varying the degree to which news are \imbued" of risk.</p>
      <p>In the current implementation, if the side e ects of new scenarios were to
be investigated, we should either build suitable synthetic documents or collect
actual documents, and provide those documents as input to the pipeline
constituted of the modules depicted in Fig. 1. However, some steps require manual
intervention, which might be cumbersome or even impossible for specialists. For
example, if a sociologist wanted to investigate how perception of risk will evolve
if news about nuclear accidents were broadcasted, s/he should wait for actual
news or generate synthetic news lled with keywords about nuclear accidents.
Actual news documents would be di cult to assess because of the assessors'
e ort required to label a training set that is large enough to train TIPS.
Although synthetic documents may in principle be generated about a topic that
is traditionally perceived as source of risk, Natural Language Processing (NLP)
technologies cannot be reliably utilized to generate these documents that a
human expert can express a genuine perception of risk when s/he is asked to assess
risk. If such technology existed, this challenge would be solved.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Contribution</title>
      <p>In this paper, we address the challenges illustrated above and emerged from the
TIPS project, i.e. the limitations that only one state can be investigated at a
time, only informative content-based evidence is utilised to implement a state,
and the impossibility of making prediction and simulation of what would happen
to states when evidence will change. To the aim of facing these challenges, we
introduce a vector-based formalism describing the main concepts of TIPS. The
formalism is given in terms of de nitions illustrated below.</p>
      <p>De nition 1 (Information object). An information object is any data
container provided with an identi er.</p>
      <p>An information object is provided with an identi er to allow applications to
connect the object, whereas search is based on the object's content. Examples of
information objects are webpages, individual images, videos, music les, query
sessions, click-through data, and other user behaviour data. In this paper, an
information object is symbolized by lower-case y and de ned as a vector of the
k-dimensional real space, since a component of y is a real number.
De nition 2 (Attribute). An attribute is what we can directly observe from
information objects such as documents or user interaction actions.
Examples of attributes are informative content descriptors (e.g. keywords and
terms), intra- and inter-object links, link anchors, meta-data, annotations, or
tweets mentioning the information object (e.g. the article). The real components
of a vector in Rd correspond to the attributes observed to measure an object.
To obtain a complete formalism, in this paper, an attribute is symbolized by
lower-case v and de ned as a vector of the k-dimensional real space, since a
component of v is a real number. For example, an attribute component may
refer to a term frequency, a click occurrence, a colour code or an encoded sound
fragment frequency. A set of independent attribute vectors form a vector basis
and therefore any linear combination thereof is a vector of the same space. For
example, the j-th basis vector of the k-dimensional space has 1 at component j
and 0 elsewhere. An attribute may occur in an information object to a certain
degree. Therefore, we have to introduce the following de nition.
De nition 3 (Weight). A weight is the degree to which an attribute is present
in an informative object.</p>
      <p>A weight is thus a real number. In particular, we formalize an attribute as a
vector of the k-dimensional real space. It is assumed an independence relationship
between the attribute vectors, so that it is possible to formalize an information
object as a linear combination of attribute vectors. If v1; : : : ; vn are n attribute
vectors, an information object vector can be written as
x = a1v1 +
+ anvn
(3)
where n is the number of attributes used to represent objects (e.g. the number
of keywords) and the a's are the attribute vector weights measuring the degree
to which an attribute describes an information object. Questions about a v can
be answered when the vectorial representation y of a user is matched against x,
thus obtaining the corresponding a.</p>
      <p>The main question is how can we model states? Contrary to attributes {
which are manifest { states can only be indirectly { they are latent { observed
from information objects.</p>
      <p>De nition 4 (State). A state is a latent characteristic of a user when
interacting with an information object.</p>
      <p>The main thrust is that a state refers to both an information object and a
user. As a state is a latent characteristic of an object-user pair, some attributes
have to be observed to make the state explicit. However rich the description of
information objects can be in terms of attributes, some hypotheses about the
user's state when s/he is interacting with objects can be explained only if the
latent, unobserved states can be explicitly modelled, thus allowing to predict and
simulate how states can evolve when the information objects { user interaction
included { are observed. In this paper, a state is symbolized by lower-case z and
de ned as a vector of the k-dimensional real space, since a component of z is
a real number. A state may occur in an information object to a certain degree.
Therefore, we have to introduce the following de nition.</p>
      <p>De nition 5 (Indicator). An indicator is the degree to which a state is present
in an informative object.</p>
      <p>An indicator is thus a real number too. As objects, states and attributes are
both placed in the same vector space, it is possible to represent an information
object as a linear combination of m state vectors as follows:
x = b1z1 +
+ bmzm
(4)
where m is the number of states, the z's are state vectors and the b's are
indicators. Eq. (4) enables to compute each indicator using linear algebra operations.
Thus, questions about a z can be answered when the vectorial representation y
of a user is matched against x, thus obtaining the corresponding b.</p>
      <p>
        The formalism de ned above stems from the theory of abstract vector spaces
widely adopted in ML and IR and speci cally in Learning to Rank (LETOR),
which is based on vector-based attribute (also known as feature) spaces and
discriminative learning [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We indeed leverage the potential of the combination
of statistical learning and IR as implemented in LETOR since this combination
has been proved to be e cient and e ective. The de nitions provided above
means that attributes, objects and states will be represented in the same vector
space. Using one single space allows us to obtain a uniform representation of
attributes, states and objects by combining them using indicators and weights
thereof, and to seamlessly apply operations on attributes, states and objects in
a similar way as suggested by di erent retrieval and learning models.
      </p>
      <p>Instead of depicting vectors using the usual arrows in a three-dimensional
plot, we exploit the visual paradigm adopted in the current implementation
of TIPS. Consider Fig. 4 which gives a pictorial description of the formalism
mentioned above. First of all, there is a temporal axis along which states evolve as
curves. Each state corresponds to a curve; \risk" corresponds to the red curve and
\con ict" corresponds to the blue curve. The y-axis refers the indicator values;
for example, the value of the indicator of con ict is 0.32 when the information
objects are those observed on June 2014. The indicator value is depicted as a
bullet placed on the curve. It is the result of a computational process that takes
attribute weights as input.</p>
      <p>Consider a scenario where a specialist user, e.g. a sociologist, working on
the e ect of S&amp;T on the society and how the society a ects the progress in
S&amp;T. Suppose that the sociologist's task is to study and comprehend the public
perception of \climate change" in the last two decades. A possible sociologist's
research question is: how have the perception of risk related to \climate change"
been changed in the last two years? In order to carry out this investigation, a
0.75
tr
o
ica 0.5
d
n
I
0.25</p>
      <p>Conflict State</p>
      <p>Risk State</p>
      <p>Conflict Indicator (z1, b1=0.32)
0 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 201 201
31 31 31 31 31 31 31 31 31 31 31 31 41 14 14 14 14 14 14 14 14 14 14 14 15 15 15 15 15 15 15 15 15 15 5 5
-01 -02 -03 -40 -50 -60 -70 -80 -90 -01 -11 -21 -10 -20 -30 -40 -50 -60 -70 -80 -90 -01 -11 -21 -10 -20 -30 -40 -50 -60 -70 -80 -90 -01 -11 -21</p>
      <p>
        Time
states indicators
attribute weights
informative objects
possible source for information are articles published in newspapers; this is, for
instance, the approach carried out in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], where the authors investigated if it is
possible to infer information about public opinion by looking at how the media
discuss controversial technoscienti c public issues. The task can be carried out by
gathering all the articles published in the \most representative" newspapers and
relevant to \climate change", manually analyse them and provide a qualitative
discussion on the change, if any, on the perception of risk when considering this
issue.
      </p>
      <p>A parallel can be drawn between the notion of state and that of category
of a classi cation system. However, state (e.g. risk) and category (e.g. nuclear
energy) are di erent each other. Although the pertinence of a news to a category
may be viewed latent, once the classi cation have been performed the fact that
an article is pertinent to the category can be used as additional attribute to
characterize the articles.</p>
      <p>In contrast, a state, e.g. risk is that it is not directly observable from a
document. Indeed, we are not considering articles about risk or on the notion of
risk, but articles that are imbued or evoke risk. It is a latent state that can be
evoked because of some words or combination of words occurring in the article
that can trigger other issues or images related to risk, or because of two di erent
viewpoints are presented with the underlying purpose of discrediting one of them.
The risk indicator is the degree to which the risk state is present in the
userobject pair or in the interaction between the user and a set of documents.</p>
      <p>Another di erence between category and state is that a state lies in user
interaction and it is not a static feature of an information object { a category
may be viewed a static feature indeed { and does not change without changing
the object content. Instead, state may change. A user who is reading news about
nuclear energy may be or not be in a risk state depending on the personal or
social context in which s/he interacting with information objects. The same
apply to states other than risk such as con ict or economic crisis.</p>
      <p>A similarity between states and categories (or classes) is multiplicity.
Similarly to the simultaneous membership of an object to di erent classes, multiple
states such as risk, con ict or economic crisis can be latent in the same
objectuser pair. Eq. (4) does indeed express the multiplicity and simultaneity of states
in object-user pairs, where the z's can be adapted to the user's interaction. One
of these z's may refer to relevance and the indicator b thereof may measure the
degree to which x is relevant. Similarly, another z may refer to risk and the
indicator b thereof may measure the degree to which x evokes risk. Thus, we
have</p>
      <p>x = brelevancezrelevance + briskzrisk
when an object x is represented in terms of latent states or
x = aa term frequencyva term frequency
+ aanother term frequencyvanother term frequency
+ aclick frequencyvclick frequency</p>
      <p>
        The multiplicity of simultaneous states in an object requires a shift from the
current state-of-the-art IR technologies based on LETOR to a novel paradigm
called LETS. When only one state { relevance is the most important one in
IR { is considered, ranking is the natural task that has to be automatized and
LETOR is an appropriate approach to relevance-based document ranking,
especially if applied to the WWW. The application to domains other than the
WWW requires to model states other than and in parallel to relevance.
Therefore, approaches other than LETOR may be useful if not necessary, since the
users might no longer be casual users and the application domain might not be
or only be about webpages to be ranked against relevance. In these contexts,
ranked document lists may no longer be the most appropriate output as
witnessed by the experience learned from TIPS. When the ranked list is no longer
the desired output to answer one single state, LETOR needs to be re ned and
the change of the output space { multiple, simultaneous states { will naturally
lead to the change of the hypothesis space { the space of functions { and of
the loss function, as well as the change of the learning theory. The new search
scenario may be decomposed into several sub ranking tasks { corresponding to
the multiple simultaneous states although many key components in learning to
rank can still be used. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] Fig. 5 depicts some di erences between the LETOR
framework and the LETS framework. The former (Fig. 5(a)) considers one main
      </p>
      <p>true
relevance
true relevance</p>
      <p>?
- relevance predict-ed</p>
      <p>state
predictor relevance
(a) LETOR framework
attribut-es
true
states
?
relevance predictor { there is one state, i.e. relevance { fueled by di erent
attributes (or features). The LETS framework (Fig. 5(b)) instead includes more
than one predictor, one predictor for each state. The prediction of these states
may be simultaneous, thus requiring parallel optimizations and predictions.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Related Work</title>
      <p>
        As the proposed project is interdisciplinary across sociology and computer
science, contributions both in computational and social sciences are relevant to it.
In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] the "six degree of separation" phenomenon was con rmed at very large
scale (Facebook Graph) and previously invisible social structures were captured.
This example can be seen as inscribed in a broad process of developing new ways
for the analysis of social phenomena built on the assumption that the Web is not
another world - the virtual one - but is a constitutive element of social reality,
and one increasingly relevant as argued in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. We use the expression SC to
refer to the research activities that involve the analysis of social phenomena in
digital resources [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]; another term is Computational Social Science.
      </p>
      <p>
        Relevant contributions in modelling informative content were proposed in IR.
In IR, Query Expansion (QE) techniques modify the initial query formulation
by extracting from documents relevant (or assumed to be relevant) to the
considered query additional descriptors to obtain a more e ective information need
representation. Many of these techniques are surveyed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that points out that
the adoption of QE on dynamic corpora is still an open issue.
      </p>
      <p>Recent works focused on temporal Web dynamics and its application to IR,
but they are mainly focused on Web user behaviour dynamics, on changing
individual document content, or on the variation of the single term collection
frequency over time. The TREC Knowledge-Based Acceleration track considers a
time ordered corpus; however, the task is ltering documents that would change
the pro le of people and organizations, and the list of entities is prede ned - this
project is not restricted to entity types.</p>
      <p>
        Topic Models (TM) aim at automatically discovering the main \themes" in
a document corpus. A well known TM is Latent Dirichlet Allocation (LDA)
where documents are modelled as a distribution over a shared set of topics,
which are themselves distributions over words generated by one of these topics.
LDA assumes a xed number of topics and the probability of seeing a topic is
independent over time - in contrast, we will also address time. These issues are
addressed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] but experiments are performed on relatively small datasets - in
contrast, we will also address scalability.
      </p>
      <p>
        This project focuses on techno-scienti c controversies, i.e. issues able to
trigger intense public debates, even in the case they address techno-scienti c
discussions such as those on climate change, GMOs, cloning, nuclear power merge
and blur the traditional boundaries among expertise, policy making, politics,
and public opinion [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Content analysis and the techniques traditionally used
by social scientists to analyse textual corpora are the basis for developing novel
indicators of techno-scienti c controversies.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Final Remarks and Future Work</title>
      <p>TIPS has been the starting point of an inter-disciplinary research project funded
with the aim of providing expert users such as sociologists and economists with
an e ective and e cient system to investigate techno-scienti c issues. Starting
from this paper, we will pursue this objective and will also make a contribution
at the level of methodology and system evaluation along two main directions.</p>
      <p>We will address the problems related to the scenario simulation mentioned in
Section 3 and will design, implement and evaluate methods for interacting with
the representation of multiple and simultaneous states, e.g. risk and con ict.
These methods will allow us to create scenarios by operating directly on state
representations, thus avoiding the need to apply the entire pipeline for each
scenario under investigation. To this end we will de ne a set of algebraic operators
and implementation thereof within the functional scheme depicted in Fig. 5(b).</p>
      <p>Besides explicitly considering interaction data when modelling states through
ML algorithms, we will integrate interaction data to provide the user with
information on the degree of a state present in the set of documents examined
in the last sessions and how this degree di ers from the \state distribution" in
the overall corpus. In other words, interaction data can signal the tendency of
a particular user to explore, say risky documents when performing a task or
accomplishing a speci c information goal. This signal can motivate the user to
explore additional parts of the informative space in order to form his/her
opinion on the issue and be \less subject" to the way the issue is presented in some
venues | e.g. a particular set of newspapers or blogs.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Alvarez</surname>
          </string-name>
          . Introduction.
          <source>In Computational Social Science</source>
          , pages
          <volume>1</volume>
          {
          <fpage>24</fpage>
          . Cambridge University Press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>U.</given-names>
            <surname>Beck</surname>
          </string-name>
          .
          <article-title>Risikogesellschaft - Auf dem Weg in eine andere Moderne</article-title>
          . Suhrkamp, Frankfurt/Main,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>C.</given-names>
            <surname>Carpineto</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Romano</surname>
          </string-name>
          .
          <article-title>A survey of automatic query expansion in information retrieval</article-title>
          .
          <source>ACM Computing Surveys</source>
          ,
          <volume>44</volume>
          (
          <issue>1</issue>
          ):1{
          <fpage>50</fpage>
          ,
          <string-name>
            <surname>Jan</surname>
          </string-name>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Dubey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hefny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Williamson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Xing</surname>
          </string-name>
          .
          <article-title>A nonparametric mixture model for topic modeling over time</article-title>
          .
          <source>In Proceedings of the 13th SIAM International Conference on Data Mining, May 2-4</source>
          ,
          <year>2013</year>
          . Austin, Texas, USA., pages
          <volume>530</volume>
          {
          <fpage>538</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kelly</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Teevan</surname>
          </string-name>
          .
          <article-title>Implicit feedback for inferring user preference: A bibliography</article-title>
          .
          <source>SIGIR Forum</source>
          ,
          <volume>37</volume>
          (
          <issue>2</issue>
          ):
          <volume>18</volume>
          {
          <fpage>28</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Lalmas</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Ruthven.</surname>
          </string-name>
          <article-title>A survey on the use of relevance feedback for information access systems</article-title>
          .
          <source>Knowledge Engineering Review</source>
          ,
          <volume>18</volume>
          (
          <issue>1</issue>
          ):
          <volume>95</volume>
          {
          <fpage>145</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>T</surname>
          </string-name>
          .-Y. Liu.
          <article-title>Learning to Rank for Information Retrieval</article-title>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Lorenzet</surname>
          </string-name>
          .
          <article-title>Fear of being irrelevant? Science communication and nanotechnology as an internal controversy</article-title>
          .
          <source>Journal of Science Communication</source>
          ,
          <volume>11</volume>
          (
          <issue>4</issue>
          ),
          <year>2012</year>
          . C04.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D.</given-names>
            <surname>Mimno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          .
          <article-title>Mixtures of Hierarchical Topics with Pachinko Allocation</article-title>
          .
          <source>In Proceedings of ICML '07</source>
          , pages
          <fpage>633</fpage>
          {
          <fpage>640</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>F.</given-names>
            <surname>Neresini</surname>
          </string-name>
          .
          <article-title>And man descended from the sheep: the public debate on cloning in the italian press</article-title>
          .
          <source>Public Understanding of Science</source>
          ,
          <volume>9</volume>
          (
          <issue>4</issue>
          ):
          <volume>359</volume>
          {
          <fpage>382</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>F.</given-names>
            <surname>Neresini</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Lorenzet</surname>
          </string-name>
          .
          <article-title>Can media monitoring be a proxy for public opinion about technoscienti c controversies? The case of the Italian public debate on nuclear power</article-title>
          .
          <source>Public Understanding of Science</source>
          ,
          <volume>25</volume>
          (
          <issue>2</issue>
          ):
          <volume>171</volume>
          {
          <fpage>185</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>R.</given-names>
            <surname>Rogers</surname>
          </string-name>
          .
          <article-title>The End of the Virtual</article-title>
          . Amsterdam University Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>R.</given-names>
            <surname>Rogers</surname>
          </string-name>
          . Digital Methods. MIT Press,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. G. Sadowsky.
          <article-title>Future developments in social science computing</article-title>
          .
          <source>In Proceedings of Spring Joint Computer Conference</source>
          , pages
          <volume>875</volume>
          {
          <fpage>883</fpage>
          ,
          <year>1972</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>A.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Buckley</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Mitra</surname>
          </string-name>
          .
          <article-title>Pivoted document length normalization</article-title>
          .
          <source>In Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <volume>21</volume>
          {
          <fpage>29</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>J. Ugander</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Karrer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Backstrom</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Marlow</surname>
          </string-name>
          .
          <article-title>The anatomy of the facebook social graph</article-title>
          .
          <source>CoRR, abs/1111.4503</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>C.</given-names>
            <surname>Warshaw</surname>
          </string-name>
          .
          <article-title>The application of big data in surveys to the study of elections, pubic opinion, and representation</article-title>
          .
          <source>In Computational Social Science, chapter 1</source>
          , pages
          <fpage>27</fpage>
          {
          <fpage>50</fpage>
          . Cambridge University Press,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>