<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hitachi Materials Informatics Analytics Platform Assisting Rapid Development</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yoshihiro Osakabe</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Akinori Asahara</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hidekazu Morita Hitachi Ltd.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marunouchi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiyoda-ku</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tokyo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>yoshihiro.osakabe.fj@hitachi.comg</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>23</fpage>
      <lpage>25</lpage>
      <abstract>
        <p>The data science platform for materials developments is demonstrated. Due to the recent great advances in artificial intelligence, it becomes more realistic that the industrial application of materials informatics (MI) which is the data-driven approach to discover and investigate materials characteristics. However, it is not quite easy for materials manufacturers to set up MI analytics environments without any help. Therefore, we provide the user-friendly cloud-based IT platform for non-experts of IT enabling materials scientists in R&amp;D departments to analyze their experimental data effectively for rapid developments.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Motivation</title>
      <p>
        Product developments require significant time and costs to
find the optimal combination of ingredients and
parameters. Materials Informatics (MI) is an emerging study field
based on the both informatics and materials science, with
the goal of greatly reducing the resources and risks required
to discover, invest, and deploy new materials
        <xref ref-type="bibr" rid="ref3">(Curtarolo et
al. 2013)</xref>
        . Recently, artificial intelligence (AI) has improved
the MI performance, thus the experimental candidates can be
narrowed down without unnecessary trials and errors before
its actual experiments to discover or create new materials
with yet-to-be realized properties. In fact, US government
has invested over $250 million to assist MI projects
        <xref ref-type="bibr" rid="ref6">(Materials Genome Initiative 2011)</xref>
        . The Novel Materials
Discovery Laboratory in EU also opens new oppotunities to
investigating MI by delivering analytics tools and open access
repository of materials data
        <xref ref-type="bibr" rid="ref7">(NOMAD Laboratory 2015)</xref>
        .
According to such outreach activities, there has been heavy
demands of materials manufacturers for introducing
MIpowered methodology into their R&amp;D processes to increase
their industrial competitiveness, and the number of startups
in MI analytics services is increasing.
      </p>
      <p>In figure 1, the concept of this demonstration is illustrated.
In many cases, it is difficult for materials scientists to select
suitable preprocessing method and effective algorithm to
solve their problems because they do not have enough
informatics knowledge, which means that they need the supports
of informatics experts (data scientists). Their relation can be
understood as that between a runner and his escort, thus this
phase can be regarded as an “accompanying phase.” Though
this service style is common, it may remain the
possibility that the informatics experts can not exactly understand
the characteristics of target materials and obtain the
knowhows materials scientists have. This problem will be solved
if materials scientists can reach analysis results by
themselves without the excessive IT and analytics knowledge.
That phase can be understood as a “self-managing phase,”
and the MI analytics services should be shifted to that phase
from accompanying phase for scaling up and rapid
prototyping. It suggests the need of the informatics expert alternative
and one-stop platform for storing, analyzing data and
visualizing analysis results.
We have developed an IT platform of MI, called Materials
Informatics Analytics Platform (MIAP), for R&amp;D teams
of various manufacturing companies. This platform brings
together all data into one place to make it easier for
researchers to access and custom machine learning algorithms
by themselves without any additional help of informatics
experts. In fact, it includes the functions that support almost
every step required for MI analytics.</p>
      <sec id="sec-1-1">
        <title>Functions</title>
        <p>MIAP is a cloud service thus the user interface is accessible
via common web browsers. It mainly includes three
functionalities; storing, analyzing and visualizing. In the
following, their details are explained.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1. Storing</title>
        <p>In this platform, all input and output data is stored in
PostgreSQL database servers. Various file types are
acceptable; CSV, Microsoft Excel, NetCDF and so on.
Graphical user interface (GUI) is utilized to upload and
import data into databases. At the same time, it also receives
SQL queries to manage data tables directly with the
implemented query editor for complecated operations. With
GUI for example, users can define the data type of each
column without typing any complicated SQL queries.</p>
      </sec>
      <sec id="sec-1-3">
        <title>2. Analyzing</title>
        <p>
          In general, MI problems are interpreted as regression and
classification tasks. Thus, it supports the various
wellknown machine learning algorithms such as Random
Forest
          <xref ref-type="bibr" rid="ref1">(Breiman 2001)</xref>
          , Gaussian Process
          <xref ref-type="bibr" rid="ref8">(Rasmussen and
Williams 2006)</xref>
          , Support Vector Machine
          <xref ref-type="bibr" rid="ref2">(Burges 1998)</xref>
          and Gradient Boosting
          <xref ref-type="bibr" rid="ref4">(Friedman 2001)</xref>
          . It also makes
predictions and optimizations possible. In addition, one
of the MIAP unique features is the implementation of the
AI-based best practices of efficient methodologies for
individual customers, which contributes to reduce their
experiment iterations. In most cases, once users have
developed their best practices, they can easily and repeatedly
apply the same method to new data by themselves.
        </p>
      </sec>
      <sec id="sec-1-4">
        <title>3. Visualizing</title>
        <p>It provides basic visualization tool to plot data in database
by selecting target column and graph types (bar, line and
pie graph). To check the learning performance, users only
have to click on automatically generated truth-prediction
scattering graphs. In addition, it provides the original UI
tool derived from a Geospatial Information System (GIS)
tool that draws animation along with time in 2D and 3D
graphs. It means that users can see the time evolution of
materials properties.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Demonstration</title>
      <p>The usage of MIAP is demonstrated by taking an example of
the search for an optimal recipe that improves the material
properties of a ready-made product.</p>
      <p>
        First, collect data accumulated in the process of making
target product. Second, upload them to the MIAP database.
MIAP automatically converts files with different formats
into a predetermined format using KNIME, the open source
software
        <xref ref-type="bibr" rid="ref5">(KNIME 2019)</xref>
        . Next, user specify the target
material property expected to be improved as an objective
variable, and other properties are set to explainable variables.
After selecting algorithm for modeling and entering the
output table name for the current attempt, the learning is started
by pressing the execute button. These operations are very
simple because almost all users have to do is just clicking on
corresponding tabs. As shown in Figure 2, users can
recognize the results are listed on results view window when the
calculation is finished. Because the automatically generated
truth-prediction scattering is shown with common indicators
to score learning performace such as Root Mean Squared
Error (RMSE) and correlation coefficient, it is possible for
users to judge whether the learning is succeeded or not. After
users can obtain well-trained model via iterational attempts,
they can predict the target material property with candidate
recipes to narrow down before the actual experiments for
new products. In this way, MIAP assists to find the optimal
recipes of ingredients or parameters, which contributes to
reduce materials development resources.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Machine learning 45(1</article-title>
          ):
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Burges</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          <year>1998</year>
          .
          <article-title>A tutorial on support vector machines for pattern recognition</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <fpage>121</fpage>
          -
          <lpage>167</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Curtarolo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Hart,
          <string-name>
            <given-names>G. L.</given-names>
            ;
            <surname>Nardelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            ;
            <surname>Mingo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Sanvito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>The high-throughput highway to computational materials design</article-title>
          .
          <source>Nature</source>
          materials
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <fpage>191</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>J. H.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Greedy function approximation: a gradient boosting machine</article-title>
          .
          <source>Annals of statistics 1189-1232.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>KNIME.</surname>
          </string-name>
          <year>2019</year>
          . KNIME, https://www.knime.com/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Materials</given-names>
            <surname>Genome Initiative</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>In Website of Materials Genome Initiative (MGI)</article-title>
          , https://obamawhitehouse.archives.gov/mgi.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>NOMAD</given-names>
            <surname>Laboratory</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>In Website of Novel Materials Discovery (NOMAD) Laboratory</article-title>
          , https://nomadcoe.eu/industry/interaction
          <article-title>-with-industry.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Rasmussen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Gaussian processes for machine learning, model selection and adaptation of hyperparameters, chapter 5</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>