<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Right to Information Query Modelling via Graded Response Model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nayantara Kotoky</string-name>
          <email>nayantara@iitg.ernet.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vijaya V Saradhi</string-name>
          <email>saradhi@iitg.ernet.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indian Institute of Technology Guwahati</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Right to Information (RTI) Act, 2005 empowers citizens of India to access information from any governmental organization. Using this Act citizens can ask questions (through RTI applications/queries) to government o ces and obtain answers. In this work we attempt to model RTI queries. Objective of modeling is to understand the latent patterns such as transparency and e ectiveness of RTI Act implementation in the RTI query-reply process which are suggestive of possible amendments in the Indian Constitution. We employ Graded Response Model (GRM, a variant of Item Response Theory) for obtaining the latent patterns. A synthetic dataset corresponding to central and state educational institutions is constructed which has close characteristics to the collected RTI query dataset. From the GRM we infer that certain institutes are highly transparent in replying to citizen's questions across various categories. We also infer that RTI act's implementation is not uniform across diverse categories within a transparent institution.</p>
      </abstract>
      <kwd-group>
        <kwd>Item Response Theory</kwd>
        <kwd>Graded Response Model</kwd>
        <kwd>Right to Information</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Right to Information (RTI) Act 2005 empowers citizens of India to access
information from any public institution (those institutions that are funded by
government). The RTI Act came into force in October 12, 2005. Through this act
citizens can inspect o cial documents, contracts, press releases, records, notes,
certi ed copies by ling an RTI application/query. Each institution appoints
a Public Information O cer (PIO) to implement the RTI act and reply to the
questions posed by citizens. Citizens submit a hard copy of their questions to the
PIO. Every RTI application costs a sum of Indian rupees ten. PIO is responsible
to reply to the query within a xed time period (typically 30 days).</p>
      <p>RTI queries form a source of information where one can witness citizens'
interaction with government establishments. Such a rich source of information
when analyzed can throw light on the sensitivities of citizens and weaknesses
in the implementations of laws. Certain acts have been amended using the RTI
statistics. In particular, RTI Act itself is amended through the following two
examples:</p>
    </sec>
    <sec id="sec-2">
      <title>1. Inclusion of Indian Postal Orders:</title>
      <p>For fee payment with an RTI application, the acceptable modes of payment
were Banker's cheque, Demand Draft or by cash. All three modes had their
own additional burdens. Both demand draft and banker's cheque had their
service charges attached, and payment by cash required visiting the public
institution in person. Indian Postal Order (IPO) is another convenient mode
of paying fees, with a nominal charge of 10%, which is Re.1 for the fee of Rs.
10. However, IPOs were not acceptable as a mode of payment, because of
which there were multiple rejections of RTI applications which were perfectly
good with their content. This event was widespread enough to catch the
government's eye, and there was discussion of its inclusion as a mode of
payment. Ultimately changes were made to the RTI Act's scope by adding
IPOs as a mode of payment. [3].
2. Exemption of political parties from being a public authority:
Asking for source of funding for political parties is not uncommon. With the
intention of understanding the inner workings of the organization, political
parties are often asked to cite their source of funding. With the advent
of RTI, multiple applications were led asking their nancial details. The
parties argued that they are not under the direct funding of the central or
state government and hence are not liable to divulge such information. Such
queries were repeatedly rejected, and a notice was issued stating political
parties as not being public authorities. It was nally included in the RTI
amendment Bill 2013 [4].</p>
      <p>From the above two examples it is observed that "repeated rejections" of
RTI queries served as a feedback for introduction of amendments into existing
RTI act. This leads us to believe that the latent patterns in the RTI query log
provide potential pointers for predicting future amendments. The objective of
this work is to collect the RTI queries and associated response (whether the
institute has replied to the query, rejected the query or referred to third party)
by institutions across India, model thus collected text data and identify latent
patterns in the RTI query database.</p>
      <p>We propose to model the RTI queries text database as a two dimensional
matrix whose rows correspond to institutions and columns correspond to topics
on which questions were posed to individual institutions. An entry ij in this
matrix correspond to percentage of replies an institute i has given against a
query topic j. This matrix is given as input to the Graded Response Model
(GRM) to identify latent patterns in the RTI query-reply process.</p>
      <p>After running GRM on our RTI data, each institution has been designated a
'transparency' value that determines how e ective an institution is with respect
to replying RTI queries, and indicates a di erence between the central and state
educational institutions. The model also identi es di erences in the query-reply
process along di erent query topics.</p>
      <p>Contributions:
1. This is the rst attempt in collecting RTI query-reply data across India.
2. A two dimensional query-reply matrix is constructed out of the given RTI
query-reply text database instead of using conventional text modeling
methods such as vector space model, latent semantic indexing, LDA etc.
3. We employ for the rst time psychometric models in RTI query text
document analysis.
2
2.1</p>
      <sec id="sec-2-1">
        <title>Related work</title>
        <sec id="sec-2-1-1">
          <title>Modelling the Political Domain</title>
          <p>Attempts to model the legislative structure and outlook have been seen in the
literature. Now and again, researchers have sought to apply mathematical
models to represent a airs in the political domain. Such work opens up scope for
understanding the political issues in depth. Gerrish and Blei [5] developed a
probabilistic model for legislative data to identify voting patterns in speci c
political issues. They used the text of bills to identify the speci c topics to which
the bills relate to, and attempted to identify what the lawmakers' stance is with
respect to bills with di erent topics (issues). They argued that a lawmaker's
attitude cannot be captured accurately on the broad political structure since they
do not exhibit enough regularity in the voting patterns. It is assumed that they
have an overall (general) political stand but have di erent political stand on
speci c issues that the bills are based on. The paper introduces an issue adjusted
model that identi es each lawmaker's position on individual topics, called the
'issue adjusted ideal-point model'. The adjusted model has been able to
identify the lawmakers' political stand in a more realistic way, and for each issue
individually.</p>
          <p>Poole and Rosenthal [6] analysed a variant of voting patterns, namely, roll
call data for legislators' votes. They took US voting data where choosers are
representatives of the law or senators, and the choices are binary, that is, yes or
no. They developed a unidimensional model of probabilistic roll call voting, and
the methods can be applied to the analysis of voting in popular elections and
other forms of political choice behavior.
2.2</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Forms of Queries</title>
          <p>From classrooms to commercial platforms and entertainment, queries are found
everywhere and in all forms. Examples include e-commerce queries, customer
service queries, product review queries, tourism queries, personal and rhetorical
queries (natural language), queries in an Issue Tracking System, queries in
medical diagnosis and of course RTI queries. Each of these query types have di erent
models of analysis. Some of the ways of modelling are:</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>1. Web search engines: Web queries (queries put to search engines for web search) are analysed to improve user experience and search engine performance. Research has</title>
      <p>been done to nd user goals from queries [7] and temporal dynamics of
query patterns have been studied [8]. [9] proposes methods for clustering
similar queries together, which helps us to understand how frequent and how
diverse web queries are. Traditional information retrieval mostly depended
on simple term matching between queries and documents. However, it has
been observed over time that understanding the meaning of the query is
important in improving the precision of a search result, like certain keywords
have more relevance in a given query and synonyms need to be identi ed.
Attempts to nd such hidden semantics in the queries have been made by
[10].
2. Question-answer (Q/A) system:</p>
      <p>Q/A systems do not retrieve documents, but give brief, relevant answers in
short text. This speciality requires time, processing power as well as
computation and understanding of the semantics of the query. In order to overcome
the bottlenecks of natural language understanding, an amalgamation of
statistical and representation based methods is required. Semantic information
in questions and answers classi cation is studied in [11]. [12] has attempted
to design a paraphrase component in a natural language question-answer
system, whereas [13] has presented a new topology to support construction
of question-answer systems.
3. Examination sets (questionnaires/test questions):</p>
      <p>Questions are used to determine the quali cation of individuals or behaviour
of events. Typical examples are the survey questions under social or
business context, tests for students, diagnosis of illness etc. Applications include
attempts to model response behaviours and nding optimum set of
questions for judgement. Examples are equating tests [14], understanding family
relationships [15] etc.
3
3.1</p>
      <sec id="sec-3-1">
        <title>Item Response Theory</title>
        <sec id="sec-3-1-1">
          <title>Description</title>
          <p>Item Response Theory (IRT) is a method for psychometric analysis. It uses
statistics to analyse how people (test takers) respond to di erent questions and
elements. The modelling of the data is done as a function that is an adjustment
between two criteria:
{ The persons abilities, perspective or personality traits and
{ The item (question) di culty.</p>
          <p>The perception behind IRT is that probability of a correct response to an
item is a mathematical function of person and the item parameters. IRT treats
di culty of each item as information to be incorporated in scaling items. The
person parameter is interpreted as a single latent trait. Example of person
parameters include intelligence, attitude etc. Likewise, we have item parameters
that are taken into consideration like di culty of the item, discrimination (slope
or correlation) representing how sharply the rate of success of persons varies
with their ability, or guessing parameter which characterises certain items that
even low intelligent persons can attempt to get correct response by guessing.
IRT has a few presumptions. The rst is that all items are independent of each
other. Hence each item is modelled separately with its own set of parameters
(which shall be discussed next). Second is that the response of a person to an
item can be modelled by a mathematical Item Response Function (IRF). Also,
the latent trait theta ( ) is assigned to each person that gives us the ability of
the persons in a unidimensional scale. The main advantage of IRT is that the
ability parameter ( ) and the item di culty parameter are modelled on the same
scale. We can imagine ability (intelligence of the student) and di culty (of the
questions) as two opposing parameters, both contributing to the probability of
keying the correct response.</p>
          <p>Measurement items with multiple response options also exist. In case of
polytomous models, each category function must be modelled explicitly. We can
imagine di erent response categories to be separated by boundaries.
Responding in a particular category means responding between two boundaries of that
response category. This gives rise to two types of conditional probabilities:
{ Probability of responding in a given category
{ Probability of responding positively rather than negatively at a given
boundary between two categories</p>
          <p>In case of polytomous items with multiple responses, in order to identify
probability of responding in a particular category, we need to identify probability
at both the boundaries. Positivity in one category boundary does not mean
response to the adjacent response category. It simply means that probability is
positive for all the subsequent categories, and might not refer to only the adjacent
response category. Hence probability of responding to a particular category shall
entail positivity at the lower category boundary and negative probability at the
upper category boundary. This idea shall be exploited in the model that we shall
use for our experiments.
3.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Graded response model</title>
          <p>The Graded Response Model (GRM) is a polytomous IRT model for ordinal
response categories. It belongs to the class of Thurston/Samejima models. The
GRM is an extension of the 2-Parameter Logic Model. Let be the latent ability
underlying the response to the test items. The probability of a candidate with
ability responding to item i in a particular category c is:
where</p>
          <p>Pic( ) = Pic( )</p>
          <p>Pic+1( )
Pic( ) =
1 + exp(
1
i(
ic))</p>
          <p>i is the Item slope parameter (one per item), ic is the Category threshold
parameters and Pic is the Category Boundary Response Function (CBRF) for
item i and category c. There is one set of i1,..., im for each item and are
ordered, where m+1 is the number of categories [16].</p>
          <p>The psychological idea behind this is that in a dataset with polytomous
response categories, each response category of an item exerts a level of attraction
on persons taking the test. In the context of an entire item, being attracted
to a category must take all prior category attractions into account. In other
words, the probability of responding in any given category is a combination of
being attracted through all previous categories up to the given category, but no
further. In the case of ordered categories, this process means that to respond in a
particular category, a person must have passed through all preceding categories.
Let Pig be the probability of responding in a particular category (g ) to item i.
If Pig represents a CBRF in the Thurstone/ Samejima models (where both are
conditional on ), then</p>
          <p>Pig = Pig</p>
          <p>Pig+1</p>
          <p>The probability of responding in a particular category is equal to the
probability of responding above (on the positive side of) the lower boundary for the
category (ig ) minus the probability of responding above the category's upper
boundary (ig+1 ).
3.3</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Parameter Estimation</title>
          <p>There are two types of parameters in IRT, that is, Item parameters and Person
(ability) parameter. Since IRT is a trade-o between the two types, both are
estimated iteratively to arrive at the best t. For polytomous data, data is
modelled by multiple dichotomizations at the category boundaries and nally using
all the information to reach the nal estimation of the parameters. For
dichotomous data, parameter estimation is done di erently for di erent parameters.
Estimating 'ability' parameter with known Item Parameters: To
estimate an examinee's unknown ability parameter, it will be assumed that the
numerical values of the parameters of the test items are known. It is an iterative
process, and begins with some known values of the item parameters. The
probability of the correct response to each item is then computed, and then the ability
parameter is slightly adjusted so that the values closely match the observed
values. The process is repeated until the adjustment becomes small enough that
the change in the estimated ability is negligible.</p>
          <p>s+1 =
s +</p>
          <p>P ai[ui Pi( s)]</p>
          <p>P ai2P ( s)Q( s)
where s is the estimated ability of the examinee within iterations, ai is
the discrimination parameter of item i, ui is the response given by examinee to
item i, Pi( s) is the probability of correct response to item i at ability and
Qi( s) = 1 Pi( s) is the probability of incorrect response to item i at ability
.</p>
          <p>Bayesian Estimation is used to estimate ability parameters given the item
parameters. We have from Bayes' theorem
which can also be written as</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Taking logarithm of both sides,</title>
      <p>f ( ju) =
f (uj )f ( )</p>
      <p>f (u)
f ( ju) / L(uj )f ( )
lnf ( ju) / lnL(uj ) + lnf ( )</p>
      <p>The posterior is directly proportional to the likelihood multiplied by prior,
where f ( j u) is the posterior estimate, L(u j ) is the likelihood estimation
and f( ) is the prior distribution. For each and every , we can calculate the
likelihood function and we also have the prior. So, we can calculate posterior
distribution P ( j u). The prior distribution has a bell shaped curve; hence the
right hand side of the equation shall have a point with slope 0.</p>
      <p>Estimating item parameter from response data: Let us divide examinees
into J groups along the scale so that all the examinees within a given group
have the same ability level j , where j = 1, 2, 3. . . . J. If rj is the examinees that
give correct response, then at an ability level of j , the observed proportion of
correct response is p( j ) = rj /mj , where mj is the total number of examinees
in the group. Now the value of rj can be obtained and p( j ) computed for each
of the j ability levels established along the ability scale. The main task now is
to nd an Item Characteristic Curve that best ts the observed proportions of
correct responses.</p>
      <p>For the estimation, initial values of item parameters are established. These
values are then used to compute p( j ) with the logistic equation. Iteratively, the
item parameters are adjusted as well to nd better values that re ect proximity
with our observed data. This process of adjusting the estimates is continued
until the adjustments get so small that little improvement in the agreement is
possible. At this point, the estimation procedure is terminated and the current
values of b and a are the item parameter estimates.</p>
      <p>The method used to calculate item parameters from response data is called
the marginal maximum Likelihood. Given the joint distribution of a function
f (x1; x2), we can calculate the marginal distribution of f (x1) as:
Z 1
f (x1; x2)dx2</p>
      <p>Let yi be the response vector for person i. Yij shall be the response given
by person i to item j. Let J be the total number of items, i be the ability of
person i and be the matrix of true item parameters. So we have,
f (yij ; ) =</p>
      <p>J
Y P yij ( i)
j=1
Hence, the marginal distribution of item parameters can be given as:
f (yij ) =</p>
      <p>Z</p>
      <p>f (yij ; )g( )d( )</p>
      <p>Let Y be the response matrix of each and every person and let there be n
persons in total. So:</p>
    </sec>
    <sec id="sec-5">
      <title>Taking logarithm of both sides for likelihood:</title>
      <p>f (Y j ) =
logL(Y j ) =
n
Y f (yij )
i=1
n
X logf (yij )
i=1</p>
      <p>The value of where likelihood function is maximised is found via Bayesian
Estimation as described above.
4
4.1</p>
      <sec id="sec-5-1">
        <title>Dataset</title>
        <sec id="sec-5-1-1">
          <title>Data Collection</title>
          <p>For the purpose of our study, we have decided to create an "RTI database" as
part of our research. Our dataset consists of the RTI applications that have been
posted to all public educational institutions by the citizens of India. The data
collected consists of RTI applications (which include the RTI queries), date of
reply of each query and the rejected queries with their grounds of rejection. This
collection is going on and the database is not yet complete.</p>
          <p>The data collection was formally started on 01.01.2015. RTI data is not
found online, but have to be collected from each individual institution. Hence
we resorted to ling an RTI application of our own asking for the data required,
namely, all the RTI applications received by the institution, date of reply of each
query and the rejected queries with their grounds of rejection. There is no facility
for online RTI ling, so we had to post our RTI application to each institution.
We started with the educational boards of high school and higher secondary
level, and moved ahead towards universities. We shall collect RTI data from a
total of 1053 educational institutions across India. Till date, we have led RTI
applications to a total of 360 institutions and have received a variety of replies
to the same application from di erent institutions, both positive and negative.
Of the institutions that received our RTI application requesting the RTI data,
56.38% have rejected our application citing various reasons. Up to this date, we
have collected data from a total of 44 institutions and 113 additional institutions
have agreed to give us the data (on payment of extra money or collect data by
visiting their o ce). The average time of receiving a reply to our application is
53.2 days. For the institutions that we have collected data from, it has taken us
an average of 73.9 days to nally receive the data. This has resulted in around
35,000 RTI applications and reply stats. Each RTI application contains multiple
queries (or sometimes just a single query). India is a multi-lingual country, and
the queries are mostly found in the local language of the area to which the
institutions belong. The data has not been processed yet, so the precise count
for total queries is unavailable.
4.2</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>Data modelling</title>
          <p>A citizen of India can ask an RTI query on any topic that is relevant to the
institution to which the application is led. There is also the provision of transfer
of the RTI application to the appropriate department if the reply or sought
document is not in the o ce that received the application. As a result, we nd
a variety of query types belonging to di erent topics. Upon closer inspection
of the data received by us, it was observed that the queries can more or less
be divided into some xed number of topics. Some topics gets more queried,
hence are popular among the masses, whereas some receive less queries. Areas
of educational institutions such as Academic (marks), research, infrastructure
are more targeted since people are more interested in knowing the workings
of these departments. Hence analysing the RTI query-reply patterns of these
speci c topics is of paramount importance.</p>
          <p>For our experiments using Graded Response Model, analysis can be done
on the reply, rejection, and appeal stats etc. This shall indicate transparency
among institutions and categories, probability of getting a query in a particular
category accepted or rejected, etc.</p>
          <p>{ Create matrices based on queries asked, queries replied, queries rejected,
queries appealed.
{ Analyse behaviour patterns of institutions in answering or rejecting queries
and identify the most frequently-asked topics.</p>
          <p>
            The GRM models items with polytomous response categories. The model
takes as input a matrix with items (questions) on one dimension and person
parameters on the other. Values are lled with the response categories for every
person to each item. Modelling consists of nding the optimum values of the
parameters of the model that best describes the data given. For our RTI data,
we can create matrices with reply stats, rejection stats and for query asked.
The matrices are lled with percentage of replied queries, rejected queries and
queries asked respectively. To draw an analogy between the two, the persons
in GRM data is represented by the institutions in our RTI data, the items are
represented by topics to which the queries belong. The response categories are
represented by percentages (
            <xref ref-type="bibr" rid="ref1 ref10 ref11 ref12 ref13 ref14 ref15 ref16 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">0-100</xref>
            ). Since the entries in the matrix represent
percentages and they are ordinal in nature, the use of GRM is appropriate.
The utility of the GRM for using on our RTI data is because of the latent
patterns that it helps to identify. With respect to our data, the ability of persons
shall represent 'transparency' of the institutions (with respect to answering or
rejecting RTI queries), and the item di culty shall denote the implementation of
the act across institutions for each topic of query. The parallelism of the latent
patterns between the typical dataset of GRM (multiple-choice questions) and
the RTI data (query-reply stats) is what makes this an interesting approach.
          </p>
          <p>Modelling each topic-wise statistics will give us a more in-depth picture of
the dynamics of the RTI query-reply process, and capture intrinsic details
hidden under an envelope of the overall performance of public institution. It is often
observed that certain sections of a public body are more e cient in its
performance while some others are lethargic. With targeted analysis of RTI queries
divided into topics, we aim to discover speci c issues or excellence regarding the
di erent divisions of the same institution.
5
5.1</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>Experiment and Results</title>
        <sec id="sec-5-2-1">
          <title>Constructing the dataset</title>
          <p>Our dataset consists of matrices constructed from the RTI database created by
us. An RTI application can have multiple queries. A survey of the data collected
has shown that queries can be more or less classi ed into some xed number
of categories or topics, each independent of each other. A few examples of such
categories are Administration, Library, Exams, Courses, Results, Academics,
Admissions, Research, and Tenders etc. Each category has its own individual
characteristic with respect to reply and rejection statistics. In order to dissect
the RTI properties and understanding the hidden traits, analysing
categorywise and institution-wise trends shall equip us with more information regarding
the implementation of the RTI act. We have created matrices with topics of
queries in one dimension and various institutions on the other. The matrices are
created such that persons (in GRM) are represented by institutions, items are
represented by topics and response categories are represented by percentages.</p>
          <p>The RTI data collection is still going on, and only a fraction of the data is
present with us. Additional tasks like translating various local languages into
English, digitizing the data that was received by us in the form of photocopies
etc. are yet to be undertaken. For the experiment, we have constructed a
synthetic matrix of reply statistics that resembles our RTI dataset (the few RTI
applications that we have collected). We have created matrices with topics of
queries in one dimension (items) and various institutions on the other (person
parameters). Matrix consists of ten institutions and ve topics. Institute 1 to
institute 6 are assumed to be central educational institutes and institute 7 to
institute 10 are state institutes. Institutions are arranged in rows, and columns
represent the ve query topics. The matrix is lled with the percentages of the
queries replied by each institution for each of the topic. The matrix with initial
values is shown in Table 1.
Table 1 gives us the raw values of our RTI dataset. In order to t this data into
GRM, the matrix needs to be modi ed. We have divided the percentages into
ve buckets as shown in Table 2. The buckets are created so that each response
category (each bucket) has a minimum amount of institutes' responses. This is
done to reduce sparsity of data by clubbing together a percentage range into a
single group. The response categories follow the Likert scale with 1 being the
lowest and 5 representing the highest rating. This is done because GRM expects
data in the form of ordinal response options. Here the ve buckets represent ve
response options and each institution responds to one of those options
corresponding to the respective reply percentages. Substituting the percentages with
the values of Table 2 results in the matrix shown in Table 3.
In order to run GRM, we have chosen the open source platform 'R'. It has a few
packages for IRT and we have used the 'ltm' package. The parameters obtained
by running GRM to our synthetic data are shown in Table 4.</p>
          <p>For each and every item, a graph is drawn between ability (latent trait) and
the probability of responding on a particular category. Such curves, called Item
Response Category Characteristic Curves for each of the ve items are shown in
Figures 1, 2, 3, 4 and 5.</p>
          <p>
            We have used Bayesian Estimate procedure for calculating the ability
parameter ( ) for each and every institute. Theta ( ) gives us the transparency of
an institution. Transparency for each institution is shown in Table 5.
The GRM has assigned an ability parameter to the ten institutions based on
the reply stats. In the context of our dataset, the ability parameter represents
the transparency of an institute. Higher the ability, more percentage of reply are
given to RTI queries, hence more transparent is an institution. From Table 5 it
is seen that Institute number 6 with the scores (
            <xref ref-type="bibr" rid="ref4 ref4 ref5 ref5 ref5">5,4,5,4,5</xref>
            ) has the highest ability
(1.623), and Institute number 9 with the scores (
            <xref ref-type="bibr" rid="ref1 ref1 ref2 ref2 ref3">1,2,2,1,3</xref>
            ) has the lowest ability
(-1.590). Arranging the institutions with respect to transparency value shows
that all central institutions except institution 2 has high transparency compared
to the state institutions.
          </p>
          <p>Each ij is the -value of transition between adjacent response categories.
It is the boundary at which the probability of the response falling in the
previous response category (left side) becomes less than 50% and the probability of
response falling on the subsequent response categories (categories on the right
side) is greater than 50%. These threshold values are di erent for di erent items,
indicating that each item is modelled di erently and the response thresholds are
not uniform across items but are dependent on the data distribution of each
item.</p>
          <p>Each and every item has a discrimination parameter. An item with high
discrimination parameter can discriminate well between institutes with high ability
and low ability. From our results, Finance has the highest discrimination
parameter and Employment has the lowest discrimination parameter. For an item
with low discrimination parameter, there is less distinction between the reply
patterns of high ability and low ability. Hence, observing reply stats of
'employment' item is not enough to decide the transparency of an institute. This
gives a sort of quality assessment for each item with respect to judging the RTI
characteristics between institutions.</p>
          <p>GRM also models probabilities of how each institution responds to di erent
items, that is, query topics. It can be observed from the results that certain
institutions (for example, institution 6 with =1.623) are very good in responding
to the nance category questions (Figure 1), but not so well in responding to
employment category questions (Figure 3). It means that a highly transparent
institution which replies e ciently to the ' nance' related queries do not reply
as e ciently to the 'employment' related queries. This reveals that there is
inconsistency in RTI reply across departments of the same institution, and leads
us to question as to why such inconsistencies are present.
6</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>Conclusion</title>
        <p>In this paper we have modelled the RTI query-reply process via Item Response
Theory (IRT). We have created a synthetic dataset that resembles our collected
RTI data in its characteristics, and tried to model it in terms of inputs to an
IRT model. We have selected GRM as the preferred model, and successfully run
it with promising results. The novelty of our approach lies in two main points.</p>
        <p>Firstly, such an analysis of RTI data has never been undertaken. We are
collecting RTI data related to each individual, from each public educational
institution and shall span multiple locations across India. Most RTI studies are
limited to speci c regions or speci c issues in that their surveys are based to
explore a xed set of problems. Our present work of applying learning algorithms
to uncover hidden traits in the RTI query-reply process is the rst of its kind.
Moreover, the application of GRM has been limited to the examination domain.
This work is a successful attempt to extend its application scope. Secondly, the
implications from the outcomes of this experiment are enormous. With this
attempt, we have assigned a transparency value to the institutions with respect
to the reply patterns of each and every institution. Our experiment with the
synthetic data reveals that the central institutions are more transparent in
replying to citizen's queries than the state institutions. A closer look into Tables
4 and 5 can help us extract further information. For example, certain
institutions (for example, institution 6) are very good in responding to the ' nance'
category questions (Figure 1), but not so well in responding to 'employment'
category questions (Figure 3). This reveals that there is inconsistency in RTI
replies across departments of the same institution, and leads us to question as
to why such inconsistencies are present. This is an indication of same laws being
implemented in di erent ways for di erent institutions as well as di erent
departments within the same institution. A solution for this may be to bring some
changes to the ordinances of the institution. Hence, this work of analysing RTI
queries and reply statistics shall also give us strong basis for proposing
amendments to the law of an institution. Once the data collection part is over, we shall
be able to apply this model to our actual RTI dataset, and the conclusions from
the results shall give us a clear picture of the laws and policies that govern our
public institutions.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. The Constitution of India, http://lawmin.nic.in/coi/coiason29july08.pdf</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <article-title>What is the Procedure of Amendment of the Constitution of India?</article-title>
          , http://www.preservearticles.com/201012251615/procedure-of
          <article-title>-amendment-ofthe-constitution-of-india</article-title>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. http://ccis.nic.in/WriteReadData/CircularPortal/D2/D02rti/10 9
          <fpage>2008</fpage>
          -
          <lpage>IR26042011</lpage>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. The Right to Information (Amendment) Bill,
          <year>2013</year>
          , http://www.prsindia.org/uploads/media/RTI%20%28A%29/RTI%20%28A%
          <fpage>29</fpage>
          %
          <fpage>20Bill</fpage>
          ,%
          <volume>202013</volume>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gerrish</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D. M.:</given-names>
          </string-name>
          <article-title>How they vote: Issue-adjusted models of legislative behavior</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          ,
          <volume>2753</volume>
          {
          <fpage>2761</fpage>
          ,
          <year>2012</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Poole</surname>
            ,
            <given-names>K. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
          </string-name>
          , H.:
          <article-title>A spatial model for legislative roll call analysis</article-title>
          .
          <source>American Journal of Political Science</source>
          ,
          <volume>357</volume>
          {
          <fpage>384</fpage>
          ,
          <year>1985</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lucchese</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orlando</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perego</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silvestri</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tolomei</surname>
          </string-name>
          , G.:
          <article-title>Discovering tasks from search engine query logs</article-title>
          .
          <source>ACM Transactions on Information Systems (TOIS)</source>
          , vol.
          <volume>31</volume>
          , no.
          <issue>3</issue>
          , 2013
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Beitzel</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <article-title>On understanding and classifying web queries</article-title>
          .
          <source>Citeseer</source>
          ,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A vector space model for automatic indexing</article-title>
          .
          <source>Communications of the ACM</source>
          , vol.
          <volume>18</volume>
          , no.
          <volume>11</volume>
          ,
          <issue>613</issue>
          {
          <fpage>620</fpage>
          ,
          <year>1975</year>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Deerwester</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumais</surname>
            ,
            <given-names>S. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furnas</surname>
            ,
            <given-names>G. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landauer</surname>
            ,
            <given-names>T. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harshman</surname>
          </string-name>
          , R.:
          <article-title>Indexing by latent semantic analysis</article-title>
          .
          <source>Journal of the American society for information science</source>
          , vol.
          <volume>41</volume>
          , no.
          <issue>6</issue>
          , 1990
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quarteroni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basili</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manandhar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Exploiting syntactic and shallow semantic kernels for question answer classi cation. Annual meetingassociation for computational linguistics</article-title>
          , vol.
          <volume>45</volume>
          , no.
          <issue>1</issue>
          , 2007
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>McKeown</surname>
            ,
            <given-names>K. R.</given-names>
          </string-name>
          :
          <article-title>Paraphrasing using given and new information in a questionanswer system</article-title>
          .
          <source>Proceedings of the 17th annual meeting on Association for Computational Linguistics</source>
          ,
          <volume>67</volume>
          {
          <fpage>72</fpage>
          ,
          <year>1979</year>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hermjakob</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravichandran</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A question/answer typology with surface text patterns</article-title>
          .
          <source>Proceedings of the second international conference on Human Language Technology Research</source>
          ,
          <volume>247</volume>
          {
          <fpage>251</fpage>
          ,
          <year>2002</year>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hermjakob</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravichandran</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Equating tests under the graded response model</article-title>
          .
          <source>Applied Psychological Measurement</source>
          , vol.
          <volume>16</volume>
          , no.
          <issue>1</issue>
          ,
          <issue>87</issue>
          {
          <fpage>96</fpage>
          ,
          <year>1992</year>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Preston</surname>
            ,
            <given-names>K. S. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parral</surname>
            ,
            <given-names>S. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gottfried</surname>
            ,
            <given-names>A. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliver</surname>
            ,
            <given-names>P. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gottfried</surname>
            ,
            <given-names>A. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ibrahim</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delany</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Applying the Nominal Response Model Within a Longitudinal Framework to Construct the Positive Family Relationships Scale</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          ,
          <year>2015</year>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Samejima</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Estimation of latent ability using a response pattern of graded scores</article-title>
          .
          <source>Psychometrika monograph supplement</source>
          ,
          <year>1969</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>