<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NIN-DSO:Neural Inertial Navigation Aided Direct Sparse Visual-Inertial Odometry</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yilin Zhao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuhang Gao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kun Wu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Long Zhao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Automation Science and Electrical Engineering, Beihang University</institution>
          ,
          <addr-line>37 Xueyuan Road, Beijing, 100191</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In indoor localization, visual-inertial odometry (VIO) is widely used due to its low cost and efective environmental perception capabilities. However, feature based VIO performs poorly in indoor pedestrian localization challenges. The feature points are prone to occlusion or tracking loss during complex pedestrian movements, and IMU is dificult to perform dead reckoning due to irregular or excessively large anomalous measurements. To address these issues, this paper proposes a VIO approach using the direct method assisted by a neural inertial network. The proposed method replaces the feature based method in front-end in VIO with the direct method to mitigate issues related to occlusion and tracking loss of feature points. It utilizes a neural inertial network to supply initial values for pixel tracking within the direct method and and integrates dead reckoning results as constraints during position optimization. Experimental results demonstrate that the method proposed in this paper exhibits higher positioning accuracy and robustness compared to existing methods in indoor pedestrian localization scenarios.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Indoor pedestrian positioning</kwd>
        <kwd>Neural inertial navigation</kwd>
        <kwd>Direct method</kwd>
        <kwd>Visual-inertial odometry</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        As a component of navigation, positioning, and timing technologies, indoor positioning technology is
widely applied in fields such as indoor robot navigation, augmented reality, the Internet of Things, and
indoor rescue. However, unlike outdoor positioning, GNSS signals are unavailable in indoor scenarios,
making it more dificult to obtain accurate and robust positioning results. Most indoor positioning
solutions typically use Bluetooth, UWB, or WiFi [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] to replace satellites for absolute positioning
which require external signals from the carrier. However, the accuracy of these methods depends on the
quality of the signal and the location of the base station. Additionally, some indoor localization methods
utilize LiDAR [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] or cameras [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ] for more accurate relative localization, among which VIO is
gaining attention and undergoing rapid development due to its afordability and efective environmental
perception capabilities.
      </p>
      <p>However, VIO does not perform optimally in indoor pedestrian localization. Most VIO systems
currently employ feature points in visual processing to recover the relative position transformation
between image frames. Nevertheless, when pedestrians carry the camera in motion, large changes in
the image field of view or occlusions can lead to feature point loss, resulting in system failure. In the
inertial component, consumer-grade IMUs are plagued by biases, noise, and thermal drift, resulting in
significant measurement errors. Additionally, due to the highly irregular nature of pedestrian movement,
position estimation based on kinematics or gait tracking is also suboptimal.</p>
      <p>
        Therefore, to address the problems in indoor pedestrian localization, we propose a VIO system assisted
by a neural inertial network to enhance the accuracy and robustness. The method employs a direct
method within the VIO framework based on DM-VIO [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which directly uses sparse pixel gray scale
changes to determine the relative position transformation between frames. Directly processing pixels
can avoid the problem of feature point loss during indoor pedestrian localization, but its initialization
also requires more accurate IMU data. On this basis, we train and deploy a neural inertial network
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to improve the accuracy of IMU-based position estimation in indoor pedestrian localization and
supplements the direct VIO with the prediction results of the network. The main contributions of this
work are as follows:
• The prediction results of the neural inertial network are utilized to assist in the initialization
between frames for the direct method, thereby improving the accuracy of inter-frame tracking.
• The proposed method is the first direct VIO aided by neural inertial navigation. The method
uses graph optimization to tightly couple direct VIO with the neural network navigation, thereby
enhancing the solution accuracy and robustness of the system.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. Visual-Inertial Odometry</title>
        <p>
          Visual-inertial navigation systems (VINS) are typically derived from visual SLAM methods integrated
with inertial sensors. The incorporation of inertial measurement data introduces stable scale information
to SLAM, thereby efectively enhancing the accuracy of image inter-frame matching. Most current
methods achieve tightly coupled VIO by fusing raw image and IMU data. The most classic approach is
the Multi-State Constraint Kalman Filter (MSCKF) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which uses an Extended Kalman Filter (EKF) to
add multiple camera poses to the state vector and employs a sliding window to maintain constraints
between image frames and the IMU for position solving. Some systems achieve data fusion through
optimization methods, with most utilizing feature point matching as constraints. The two most classic
methods are ORB-SLAM [
          <xref ref-type="bibr" rid="ref10 ref6">6, 10</xref>
          ] and VINS-Mono [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. ORB-SLAM2 [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] uses FAST corners and BRIEF
descriptors to form ORB features, implementing monocular VO by using feature point reprojection error
as the cost function. In VINS-Mono [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the algorithm employs Harris corner points as feature points
and utilizes the L-K optical flow method to track the feature points, establishing visual part constraints.
Additionally, the algorithm adopts a marginalization strategy that introduces a priori marginalization
error to improve the computational eficiency and robustness.
        </p>
        <p>
          With the improvement in computational performance, VO methods can process more pixel data in
realtime, leading to the emergence of direct VIO methods, which difer from feature point based approaches.
Direct methods operate directly on the image pixels captured by the camera, thus eliminating the
need for stable extraction and matching of feature points in the environment. The Direct Sparse
Odometry (DSO) algorithm [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] integrates pixel association, photometric error optimization, position
estimation, and sparse point cloud generation into a unified nonlinear optimization problem, discarding
the traditional separation between front-end and back-end in VO. Building on this, the VI-DSO [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
algorithm combines the DSO with IMU data, while DM-VIO [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] incorporates a delayed marginalization
strategy into VI-DSO to enhance the robustness and accuracy of position estimation in direct method
based VIO. However, since these methods track pixels solely through photometric invariance, they
require better initial values of relative position to achieve stable and accurate position estimation.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Neural Inertial Navigation</title>
        <p>
          The emergence of neural inertial network methods enables deep analysis of IMU measurement data,
extracting underlying motion patterns in complex movements. The RIDI [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] algorithm uses a Support
Vector Machine (SVM) model to classify the motion states of pedestrians and a Support Vector Regression
(SVR) model to regress the IMU data, thereby correcting low-frequency errors in IMU measurements.
IONet [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] treats inertial localization as a time series learning problem, utilizing Long Short Term
Memory (LSTM) to process inertial data within a time window to estimate the carrier’s position and
trajectory in a polar coordinate system. Ronin [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], based on ResNet, LSTM, and Temporal Convolutional
Network (TCN) architectures, directly estimates positions through a neural inertial network. However,
the relative position estimation of these networks lacks precise, necessitating the addition of accurate
constraints.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Neural Inertial Navigation Aided VIO</title>
        <p>In VIO methods, the processing of IMU data still adopts the method of integration using kinematic
models in traditional navigation methods. In order to decrease the computational load and efectively
improve the accuracy of inter-frame relative position estimation, the integration is typically transformed
into an inter-frame pre-integration that remains unafected by the initial state. However, these kinematic
models perform poorly in complex motion scenarios, as irregular or excessive IMU measurements can
cause the system state to gradually diverge.</p>
        <p>
          Some algorithms have incorporated position prediction results from inertial neural networks into
combined navigation systems, but fewer of them add inertial neural networks to VIO. TLIO [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] uses a
neural network to learn the 3D displacement transformation and covariance directly from a time series
of IMU data, then applies the results obtained from the network predictions to a Kalman filter to correct
state quantities such as direction, velocity, position, and IMU deviation. RNIN-VIO [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] firstly utilizes a
neural inertial navigation to assist feature point based VIO, incorporating network predictions into the
optimization model in the back-end, to enhance the accuracy and robustness of position estimation.
        </p>
        <p>However, it is worth noting that all of the above methods only involve the neural network inference
results as additional constraints added to the feature point based VIO. This structure leads to neural
network inference results with less impact on such systems. In contrast, the direct methods, although
more accurate, require higher IMU data quality. Bad IMU data will cause the direct method to fail to
initialize properly and have poorer results, which has a stronger dependence on IMU data.</p>
        <p>Therefore, we implement a neural inertial navigation aided method VIO. The proposed method is
the first method to combine inertial neural networks with direct method VIO, and it combines inertial
neural networks more tightly than the existing methods. The method efectively improves the accuracy
and robustness of the direct VIO initialization and solution.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <sec id="sec-3-1">
        <title>3.1. Overview</title>
        <p>
          boxes in Fig. 1.
The proposed method in this paper is a direct visual-inertial odometry, which improves upon the
DM-VIO framework, as illustrated in Fig. 1. After the visual part completes initialization and the
neural inertial navigation accumulates suficient IMU measurement data, the network will estimate
the position of carrier. The method utilizes the network’s predicion results to set an initial relative
position between image frames, providing a better initial photometric error for the visual part. In the
joint optimization phase, the prediction results are also integrated into the factor graph as a constraint
for position estimation. Compared to DM-VIO, the improvements of our method are highlighted in red
In direct VO [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], the initial relative position is determined by kinematic modeling. Specifically, various
photometric errors are calculated based on assumptions such as stationary, constant velocity, or constant
acceleration motions. The minimum among these errors is selected as the initial value for iterative
photometric error minimization, with the corresponding relative position used as the initial position.
        </p>
        <p>
          In DM-VIO [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], the introduction of IMU measurements enables the derivation of a coarse relative
position estimate through pre-integration. The relationship serves as the initial value, and the photometric
error calculated from this value is used as the initial error for iteration, as shown in (1).
 =
∑︁ ||( [︀ ′ −  ) − ( [] − )  /
︀]
||
(1)
where  represents a small neighborhood around point .  is the image frame.  and  are the
sequence numbers of the image frames.  represents the exposure time.  and  are the factors for
(2)
(3)
of the photometric error.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Neural Inertial Network</title>
        <p>
          We utilize the neural inertial network and proposed in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] to supplement the visual-inertial odometry.
And the data used for network training was also taken from [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The network comprises a ResNet18
structure, an LSTM structure, and fully connected layers in a cascade. The ResNet18 structure primarily
facilitates supervised learning of pedestrian motion patterns and estimation of position changes. The
LSTM structure analyzes the hidden states of pedestrian over a period of time, which may implicitly
influence the current state, to avoid the divergence problem in dead reckoning due to abnormal
IMU measurements. During the network operation, the IMU data must be transformed from the
IMU coordinate system to the VIO coordinate system, with gravity and bias efects removed. After
preprocessing the IMU measurements, they serve as the input to the network, enabling the direct
acquisition of relative position estimation of pedestrian and their corresponding covariance, as shown
in (3).
correcting the afine brightness transformation.
projected point obtained by pre-integration results.
        </p>
        <p>is the weight related to the gradient. ′ is the</p>
        <p>DM-VIO incorporates photometric errors and IMU residuals for optimization process. And its energy
function is</p>
        <p>=  ℎ + 
which consists of the photometric error ℎ, pre-integration residuals . And  is the weight
︁(
∆ 
′ ,  
′ )︁
= (︀ (−  , −  ), . . . , (, ), ℎ−</p>
        <p>︀)
 = ( − ) − 
 = ( − )
∆ ′ and its covariance  ′ are the outputs of the network.
where  (∙ ) is the function obtained by neural network fitting.  and  are the raw acceleration and
angular velocity measured by the IMU sensors.  and  are the bias obtained from the VIO system.
 is the gravity vector.  is the VIO coordinate system while  is the IMU coordinate system. 
and  are the acceleration and angular velocity at the th IMU time step in VIO coordinate system,
respectively. ℎ is the hidden state produced by LSTM at the last time step. The relative displacement
3.4. NIN-DSO</p>
        <sec id="sec-3-2-1">
          <title>3.4.1. Time Synchronization</title>
          <p>To ensure that the predicion results of network facilitate efective convergence of the VIO position
estimation, we bundle these predictions to the image inputs for approximate temporal synchronization.
Indeed, predictions from the neural inertial network resemble IMU pre-integration, serving as a type of
constraint between images. Therefore, the proposed method adjusts the prediction output frequency
to match the images input frequency. Additionally, in order to minimize the temporal discrepancy
between the images and the predition results, we interpolate the network inference results at the image
input moments. The temporal relationship of IMU data, image data, and network prediction results is
illustrated in Fig. 2.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.4.2. Coarse Tracking</title>
          <p>as shown in (4).</p>
          <p>In this paper, we use the relative position transformation predicted by neural inertial networks as an
alternative for initializing the photometric error, providing greater robustness than IMU pre-integration,
 =
∑︁ ||( [︀ ′ −  ) − ( [] − )  /
︀]
||
∈
items are consistent with (1).
where ′ is the projected point obtained by prediction results of network. And the remaining</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3.4.3. Visual-Inertial Optimization</title>
          <p>The VIO proposed in this paper achieves position estimation by minimizing the energy function ,
which comprises photometric error, pre-integration residuals, and network prediction position residuals.
Unlike the feature point based method, our approach simultaneously optimizes the position, sensor
error, scale and 3D structure, resulting in higher accuracy. The factor graph for the proposed method is
shown in Fig. 3.</p>
          <p>Based on DM-VIO, the method further incorporates the position predictions from the neural inertial
network as additional constraints. And its energy function is</p>
          <p>=  ℎ +  + 
which consists network prediction position residuals . The photometric error is the sum of the
individual pixel photometric erros, and can be determined by (6).</p>
          <p>ℎ = ∑︁ ∑︁
∑︁</p>
          <p>∈ ∈ ∈()
where  is the set of keyframes .  is the set of points in keyframes. And () is the set of points
that are jointly observed in keyframes.
(4)
(5)
(6)
︁(  ⊙ ˆ ︁) 
*
∑︁ − 1
 , *
︁(  ⊙ ˆ
︁)
corresponding covariance of the state. And ⊙ is the subtraction operation in Lie algebra.
where  is the state obtained from pre-integration. ˆ is the estimated state of VIO. ∑︀ , is the</p>
          <p>
            The IMU measurements are preprocessed before being input into the network in our method. As a
result, the relative position transformation predicted by the network remains consistent with the VIO
coordinate system. Based on this characteristic, in the part of neural inertial navigation, we use the
predictions as direct constraints between image frames rather than the constraints in [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ], as shown in
(8).
          </p>
          <p>=︁( 
 ⊙
ˆ ︁) 
*
∑︁ − 1
, *</p>
          <p>⊙
︁( 
ˆ
︁)
corresponding covariance of the state.
where</p>
          <p>is the state obtained from Neural inertial network prediction results. ∑︀, is the</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments and Results</title>
      <p>This section presents the test results on various datasets to demonstrate the efectiveness of our method.
with Ubuntu 20.04 and NVIDIA GeForce RTX3060 12GB graphics card.</p>
      <p>All the experiments were conducted on an Intel Core I5-10510U CPU@1.80GHz× 8 computer equipped</p>
      <sec id="sec-4-1">
        <title>4.1. TUM-VI Dataset Experiments</title>
        <p>
          In this subsection, we evaluate the performance of the proposed method using the TUM-VI dataset
[
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The dataset is a widely used public dataset collected by a pedestrian. However, since only the
room scenario ofers full ground truth, we present only the experimental results for this data sequence.
The absolute trajectory error (ATE) in this paper were obtained from the EVO toolbox [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], but the
APE values are larger than ATE reported in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] because we did not use EVO to fix the scale of the
algorithm’s results. The comparison of ATE between VINS-Mono, NIN-VINS [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], DM-VIO, and our
proposed method NIN-DSO is shown in TABLE 1 below.
        </p>
        <p>From the data in TABLE 1, it is evident that the position estimation of direct method based VIO
outperform feature point based method, regardless of whether an neural inertial network is added. This
indicates that direct method VIO is more suitable for indoor pedestrian positioning scenarios. Our
method achieved the best results across all data in the room sequence, reducing the average ATE by 60.2%
compared to DM-VIO. Notably, in the fourth data set, the ATE was reduced by 84.2%, demonstrating
that the addition of neural network predictions efectively addresses the issue of DM-VIO’s inability to
track and converge efectively when IMU data quality is poor.
(7)
(8)</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Real Experiment</title>
        <p>Compared to the room sequence of the TUM-VI dataset, indoor pedestrian positioning typically involves
faster movement speeds, more complex lighting environments, and larger movement scales. Additionally,
VIO tends to accumulate significant errors over long periods of operation. Therefore, we collected a
custom dataset to supplement the experiments. This dataset was collected using a Flir camera and
a Livox Avia LiDAR as the acquisition sensors, with IMU measurements provided by the Avia. The
frequency of image data is 10Hz and the resolution is 1440× 1080. The frequency of IMU data is 200Hz.
And the two sensors are synchronized by hardware trigger. The pseudo ground truth was obtained
from a higher precision Lidar-Visual-Inertial Navigation System to evaluate the performance of the
method. In this dataset, we carried the device around the interior of the New Main Building at Beihang
University. The total distance covered is 513.56 meters, and the duration of the data collection lasts
368.41 seconds. The data contains typical pedestrian positioning characteristics, and an example of
large attitude change and complex lighting environment from the data is shown in Fig. 4.</p>
        <p>As with the experiments in the first subsection, the performance of the algorithms is still evaluated
using ATE. The comparison of the results among VINS-Mono, NIN-VINS, DM-VIO, and NIN-DSO are
shown in TABLE 2.</p>
        <p>From the TABLE 2, it can be seen that our method still performs well on this custom data, achieving a
50.7% reduction in ATE relative to the best method among the other three. The same conclusion can be
drawn from experiments on the public dataset. Direct methods are more suitable for indoor pedestrian
localization than feature point based methods. Additionally, incorporating prediction of neural inertial
network can efectively address issues where the system fails to track and converge properly, mitigate
cumulative errors during long term positioning, and consequently improve the accuracy and robustness
of VIO.</p>
        <p>The estimated trajectories and ATE curves for diferent method on this dataset are shown in Fig. 5,
wiht the starting points of these trajectories marked with triangles. As can be seen in Fig. 5, the accuracy
using the direct method is higher than the feature point method. In the case of starting positions with
similar errors, our method has better scaling in the middle of the trajectory compared to DM-VIO.</p>
        <p>(a) The trajectories of diferent method.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>(b) The ATE curve of diferent method.</p>
      <p>This paper proposes a direct visual-inertial odometry with the assistance of a neural inertial network.
The neural network is implemented and trained using PyTorch, and deployed within the VIO framework
utilizing ONNX and TensorRT. The inclusion of neural inertial navigation provides better initial values
for the image inter-frame tracking in direct method based VIO, leading to a more stable visual
initialization process than DSO and DM-VIO. Furthermore, by incorporating the network’s position prediction
results into the energy function of the optimization model, in addition to photometric error and IMU
residuals, the system reduces its reliance on the visual component in complex motion environments
and mitigates errors and divergence of IMU data. Experimental results demonstrate that the proposed
neural inertial navigation aided direct VIO achieves more accurate and robust position estimation in
indoor pedestrian movement scenarios.</p>
      <p>Although the proposed method does not require stable feature point extraction and matching as
feature based methods, it relies on visual component as the foundation for system calculations, rendering
it unable to estimate positions when the visual component fails. The current supplementation using
the neural inertial network does not fundamentally solve the problem. We plan to further leverage
the powerful data processing capabilities of neural networks to evaluate visual inter-frame position
estimation results using IMU data. By using this approach, we aim to reduce the reliance of the direct
VIO on the visual component, enabling the system to continue functioning normally even when the
visual component fails.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The project is supported by the National Natural Science Foundation of China (Grant No. 42274037), the
Aeronautical Science Foundation of China (Grant No. 2022Z022051001), and the National Key Research
and Development Program of China (Grant No. 2020YFB0505804).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Cbwf: A lightweight circular-boundary-based wifi fingerprinting localization system</article-title>
          ,
          <source>IEEE Internet of Things Journal</source>
          <volume>11</volume>
          (
          <year>2024</year>
          )
          <fpage>11508</fpage>
          -
          <lpage>11523</lpage>
          . doi:
          <volume>10</volume>
          .1109/JIOT.
          <year>2023</year>
          .
          <volume>3329825</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>A Novel System for WiFi Radio Map Automatic Adaptation</article-title>
          and
          <source>Indoor Positioning</source>
          <volume>67</volume>
          (
          <year>2018</year>
          )
          <fpage>10683</fpage>
          -
          <lpage>10692</lpage>
          . doi:
          <volume>10</volume>
          .1109/TVT.
          <year>2018</year>
          .
          <volume>2867065</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Shan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Englot</surname>
          </string-name>
          ,
          <article-title>Lego-loam: Lightweight and ground-optimized lidar odometry and mapping on variable terrain</article-title>
          ,
          <source>in: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>4758</fpage>
          -
          <lpage>4765</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>Fast-lio: A fast, robust lidar-inertial odometry package by tightly-coupled iterated kalman filter</article-title>
          ,
          <source>IEEE Robotics and Automation Letters</source>
          <volume>6</volume>
          (
          <year>2021</year>
          )
          <fpage>3317</fpage>
          -
          <lpage>3324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <article-title>Vins-mono: A robust and versatile monocular visual-inertial state estimator</article-title>
          ,
          <source>IEEE Transactions on Robotics</source>
          <volume>34</volume>
          (
          <year>2018</year>
          )
          <fpage>1004</fpage>
          -
          <lpage>1020</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mur-Artal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Tardós</surname>
          </string-name>
          , Orb-slam2:
          <article-title>An open-source slam system for monocular, stereo, and rgb-d cameras</article-title>
          ,
          <source>IEEE transactions on robotics 33</source>
          (
          <year>2017</year>
          )
          <fpage>1255</fpage>
          -
          <lpage>1262</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Von Stumberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cremers</surname>
          </string-name>
          ,
          <article-title>Dm-vio: Delayed marginalization visual-inertial odometry</article-title>
          ,
          <source>IEEE Robotics and Automation Letters</source>
          <volume>7</volume>
          (
          <year>2022</year>
          )
          <fpage>1408</fpage>
          -
          <lpage>1415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Herath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Furukawa</surname>
          </string-name>
          , Ronin:
          <article-title>Robust neural inertial navigation in the wild: Benchmark, evaluations</article-title>
          , &amp;
          <article-title>new methods</article-title>
          ,
          <source>in: 2020 IEEE international conference on robotics and automation (ICRA)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>3146</fpage>
          -
          <lpage>3152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Mourikis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. I.</given-names>
            <surname>Roumeliotis</surname>
          </string-name>
          ,
          <article-title>A multi-state constraint kalman filter for vision-aided inertial navigation</article-title>
          ,
          <source>in: Proceedings 2007 IEEE international conference on robotics and automation, IEEE</source>
          ,
          <year>2007</year>
          , pp.
          <fpage>3565</fpage>
          -
          <lpage>3572</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Campos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Elvira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J. G.</given-names>
            <surname>Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Montiel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Tardós</surname>
          </string-name>
          , Orb-slam3:
          <article-title>An accurate open-source library for visual, visual-inertial, and multimap slam</article-title>
          ,
          <source>IEEE Transactions on Robotics</source>
          <volume>37</volume>
          (
          <year>2021</year>
          )
          <fpage>1874</fpage>
          -
          <lpage>1890</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Engel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Koltun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cremers</surname>
          </string-name>
          ,
          <article-title>Direct sparse odometry</article-title>
          ,
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>40</volume>
          (
          <year>2017</year>
          )
          <fpage>611</fpage>
          -
          <lpage>625</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Von Stumberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Usenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cremers</surname>
          </string-name>
          ,
          <article-title>Direct sparse visual-inertial odometry using dynamic marginalization</article-title>
          ,
          <source>in: 2018 IEEE International Conference on Robotics and Automation (ICRA)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>2510</fpage>
          -
          <lpage>2517</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Shan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Furukawa</surname>
          </string-name>
          , Ridi:
          <article-title>Robust imu double integration</article-title>
          ,
          <source>in: Proceedings of the European conference on computer vision (ECCV)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>621</fpage>
          -
          <lpage>636</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Markham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Trigoni</surname>
          </string-name>
          ,
          <article-title>Ionet: Learning to cure the curse of drift in inertial odometry</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>32</volume>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Caruso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ilg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Mourikis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Daniilidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Engel</surname>
          </string-name>
          , Tlio:
          <article-title>Tight learned inertial odometry</article-title>
          ,
          <source>IEEE Robotics and Automation Letters</source>
          <volume>5</volume>
          (
          <year>2020</year>
          )
          <fpage>5653</fpage>
          -
          <lpage>5660</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bao</surname>
          </string-name>
          , G. Zhang, Rnin-vio:
          <article-title>Robust neural inertial navigation aided visual-inertial odometry in challenging scenes</article-title>
          ,
          <source>in: 2021 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>275</fpage>
          -
          <lpage>283</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Schubert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Goll</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Demmel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Usenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stuckler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cremers</surname>
          </string-name>
          ,
          <article-title>The tum vi benchmark for evaluating visual-inertial odometry</article-title>
          ,
          <source>in: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>1680</fpage>
          -
          <lpage>1687</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Grupp</surname>
          </string-name>
          ,
          <article-title>evo: Python package for the evaluation of odometry and slam</article-title>
          ., https://github.com/ MichaelGrupp/evo,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Nin-vins: Neural inertial navigation aided visual-inertial system for pedestrian dead reckoning</article-title>
          , in: International Conference on Guidance,
          <source>Navigation and Control</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>6868</fpage>
          -
          <lpage>6878</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>