<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Hybrid Model For Information Retrieval From Microblogs During Disaster</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Du Xin</string-name>
          <email>duxin111@outlook.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qi Limin</string-name>
          <email>qilimin111@outlook.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Science, Agriculture and Engineering, University of Newcastle upon Tyne</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Remove @username</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computer Science and Technology, Heilongjiang Institute of Technology</institution>
          ,
          <addr-line>Harbin</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>School of Physical Sciences, Harbin Normal University</institution>
          ,
          <addr-line>Harbin</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>When the disaster occurs, the social such as Twitter is increasingly being helping direct rescue operati.onTshis article describes the methods we used Firien201t7h.e We regarded tdhiestinction onfeed-tweets and availability-tweets as classicfaition tasks, and the logistic regressionand Support Vector Machine are used to decidetypethe of the tw.eeItns the need and availability matching, we regard information retrievaltask, using thretrieval model to complete the task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>2. Task 1
2.1 Problem Description
2.2 Classifier Models</p>
      <p>For Task w1,e need to judge
availability-tweets and othe.rs We
classification problem.We use the
distinguish the useful -tnweedets and the</p>
      <p>th-etweentese,d
deem it as
classifiers
availa-bility
to</p>
      <p>In this working note, two
a is a binary model, which
significant liner classifier in
learning strategy is the interval
groups adopt SVM, which
is defined as the most
the feature space. The
maximization. Solving
the problem into Convex dratQicua Programming Remove URL
Problem. We chose the RBF kernel, which can map the
samples no-nlinearly to a higher dimension. The RBFID
core has fewer numerical complexity characteristics.</p>
      <p>Ultimately two groups completed Task 1 with LibSVMR.emove stop word</p>
      <p>The third group uses LRl. mEaocdhe data point of
the LR model has an impact on the classification plane.</p>
      <p>Its influence is far from its distance to the classificRaetimonove punctuation
plane, and if the data dimension is high, the LR Rmemodoevle URL
will match the parameters regularization method. In
classification, the computation is minimal, and the Note:
speed is breakneck, so the storage resources are low, so
it is convenient to observe the probability scores o(1f,
samples.
2.3 Feature Selection</p>
      <p>For feature selection, we conduct experiments with
words and -grnam (n = 2, 3, 4, 5, respectively. Th(e4,
results show that we have the best effect of using
as features. Table 1 shows the comparison of the
results of each feature. For the sipnrgeprocfes tweet
text, each team adopts different methods to preprocess
the data. Table 2 specifically describes the differ3en.tTask 2
teams in the -treparetment differences exis(t√.
indicates that this method is used, and × is not 3u.s1edP,roblem Description
w(5o,rdsHLJIT201-7IRMIDIS_1_task2_2
test
(6, HLJIT201-7IRMIDIS_1_task2_3</p>
      <p>HLJIT201-7IRMIDIS_1_task2_1
√
2_1
√
√
√
√
√
2_2
√
√
√
√
=
=
=
=
=
=
√
2_3
×
√
√
√
1_1
1_2
1_3
2_1
2_2
2_3</p>
      <p>HLJIT201-7IRMIDIS_1_task1_1
(2,</p>
      <p>HLJIT201-7IRMIDIS_1_task1_2
(3, HLJIT2017-IRMIDIS_1_task1_3</p>
    </sec>
    <sec id="sec-2">
      <title>Note:</title>
      <sec id="sec-2-1">
        <title>2-gram</title>
      </sec>
      <sec id="sec-2-2">
        <title>3-gram</title>
      </sec>
      <sec id="sec-2-3">
        <title>4-gram</title>
      </sec>
      <sec id="sec-2-4">
        <title>5-gram</title>
        <p>word</p>
      </sec>
      <sec id="sec-2-5">
        <title>Remove punctuation</title>
        <p>Pre
0.281
0.482
0.428
0.385
0.517</p>
        <p>Re
0.245
0.254
0.245
0.2
0.281
1_1
√
√
√</p>
        <p>Pre
0.477
0.671
0.641
0.604
0.709
1_2
√
√
√</p>
        <p>Re
0.474
0.597
0.606
0.556
0.644
1_3
×
√
√</p>
        <p>HLJIT2017-IRMIDIS_1_task1_1 and HLJIT20-17
IRMIDIS_1_task1_2 is different in stop words list.</p>
        <p>Task 2 requires that th-teweentesed match in Task
1 be searched by 1. TaNskee-dtweets is used as a query
set Q. Availabi-ltiwtyeets can be used as a collection of
documents D. We use statistical language models to
solve the problem of Task 2. Language models are used
to assess what kind of word sequences are more typical
according to language usages, and inject the right bias
accordingly into the system to prefer an output
sequence of words with high probability according to
the language model. If a document language model
gives the query a high probability, the query words
must have top opportunities according to the document
language model, which further means that the query
words frequently occur in the document.</p>
        <sec id="sec-2-5-1">
          <title>3.2 Relation</title>
          <p>The correlation calculation can be expressed
as shown Fingure ,1 using the -NTeweedets (Na,s the
query set Q, the Avai-ltawbeileitsy (A, as
document set D, and then the correlation
obtain the correlation R (Q, D,.</p>
          <p>briefly
the
calculation to
N</p>
          <p>Q</p>
          <p>D</p>
          <p>A
3.3 Language Model
4. Experimental Result</p>
        </sec>
        <sec id="sec-2-5-2">
          <title>4.1 Data Set</title>
          <p>At
released
tweets
Later, a
the
and</p>
          <p>start theof track,
(training
set,,
along
about
with
20,000</p>
          <p>tweets
a
samp-le
of
availabi-ltiwtyeets
in
these K20 tweets.
new
set
of
50,000 tweets released
(test set,.
4.2 Evaluation index
number of extracted
messages</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Precision</title>
      <p>=
number
of correct
mesrsaacgteesd e/xt
Recall = the
number of correct
messages extracted
the
number of
messages in the sample
 − 
=
( 2 + 1)
 2 + 
β is the
parameteri,s the
precision,  ainsd the
recall.</p>
      <p>Map：
 ̂
= 
 ̂
=</p>
      <p>( | ) ( )
∫⊖  ( | ′) ( ′)  ′
probability</p>
      <p>of the
provide
does
extra</p>
      <p>value
not appear in
the
SVM
LR
word
to
Task
precision</p>
      <p>2
but
other</p>
      <p>the
niened Task
threshold.</p>
      <sec id="sec-3-1">
        <title>Avail</title>
        <p>/
ability
ID</p>
      </sec>
      <sec id="sec-3-2">
        <title>Precision @100</title>
      </sec>
      <sec id="sec-3-3">
        <title>Recall</title>
        <p>Map</p>
      </sec>
      <sec id="sec-3-4">
        <title>Recall @1000 MAP</title>
      </sec>
      <sec id="sec-3-5">
        <title>Need</title>
      </sec>
      <sec id="sec-3-6">
        <title>Precision</title>
        <p>Average MAP
1_1
0.550
1_2
0.100
1_3
0.760
4.4 Experimental results
2_1
2_2
2_3</p>
      </sec>
      <sec id="sec-3-7">
        <title>Precision@5 0.088 0.088 0.082</title>
        <p>0.021
0.217
0.147</p>
      </sec>
      <sec id="sec-3-8">
        <title>F-score 0.034 0.125 0.105</title>
        <sec id="sec-3-8-1">
          <title>5. Conclusion</title>
          <p>Through the above experimerenstualts,
that the classification and sorting
some informal occasions are very
the future experiments, we will
machine learning and try to
features, to filter andose chtoext
occasions.</p>
          <p>we found
of text content for
regarding words. In
deepen the study of
select more different
content of informal
Acknowledgments
This work is supported by Philosophy and Social
Science Planning Project of Heilongjiang
Province, China (No. 16EDD05)</p>
          <p>A.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Overview of Retrieval from (IRMiDis,</article-title>
          . In for Information India, December-
          <volume>108</volume>
          , Ghosh,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          and
          <string-name>
            <surname>M.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>the FI2R0E17 track: Information Microblogs during Disasters Working notes of FIR-EForu2m017 Retrieval Evaluation</article-title>
          , Bangalore,
          <year>2017</year>
          ,
          <string-name>
            <given-names>CEUR</given-names>
            <surname>Choudhury</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          , Making -lSacrgalee
          <source>SVM Learning Practical. Advances in Kernel Me-thSoudpsport Vector Learning</source>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schölkopf</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Burges</surname>
          </string-name>
          and Smola (ed.,,
          <string-name>
            <surname>M-PITress</surname>
          </string-name>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[3] QI Haoliang1,CHENG Xiaolong1, YANG Muyun2, HE Xiaoning3, LI Sheng2, LEI Guohua1. High Performance Chinese Spam Filter</article-title>
          .
          <year>2010</year>
          , :
          <fpage>274</fpage>
          -
          <lpage>6</lpage>
          (
          <issue>824</issue>
          ,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>