<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>from Pyrolysis of Biomass: A Comparative Study of Imputation Algorithms and Model Benchmarking</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Antonio Elia Pascarella</string-name>
          <email>antonioelia.pascarella@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Rome, Italy</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>21</institution>
          ,
          <addr-line>80125 Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Electrical Engineering and Information Technology (DIETI), University of Naples Federico II</institution>
          ,
          <addr-line>Via Claudio</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The exhaustion of non-renewable fossil fuels has heightened awareness about environmental issues. As a result, biomass energy has come into the spotlight as a promising renewable alternative, particularly in the context of bio-oil production through pyrolysis from waste biomasses. Unfortunately, physics-aware models pose dificulties when modelling bio-oil production, prompting researchers to lean towards data-centric approaches. To cope with this problem, this paper showcases a comprehensive dataset of nearly one thousand records sourced from prior literature about bio-oil production. Besides collecting, cleaning and organising the gathered data, we also used machine learning techniques to evaluate the resulting dataset, with the most promising result yielding a mean absolute error of 2.6 and an adjusted R-squared of 0.9 in prediction of bio-oil yield. To the best of our knowledge, this paper delivers introduces the most comprehensive dataset ever collated in this domain. The assembly of such an exhaustive dataset is pivotal for sustainable process engineering because it fosters precise modelling, thus better-addressing uncertainties inherent in the process.</p>
      </abstract>
      <kwd-group>
        <kwd>Model Benchmarking</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Missing data imputation</kwd>
        <kwd>Renewable energy</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        The accelerated depletion of non-renewable fossil fuel reserves has led to critical issues related to
environmental degradation. Consequently, biomass energy, which is abundant and sustainable,
has attracted considerable attention [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Among the various conversion techniques for biomass,
the production of bio-oil through pyrolysis, and particularly its modeling, is the focus of this
work. If appropriately upgraded, bio-oil could serve as an alternative to fossil fuels.
      </p>
      <p>
        The uncertainty in the yield of bio-oil production and its qualities and the optimization of
process variables can be managed with mathematical models. Simulating pyrolysis for bio-oil
production using physics-based models is challenging. As artificial intelligence, and in particular,
the subfield of machine learning evolves, many data-centric approaches have been proposed
for estimating bio-oil yields from waste biomass properties and plant operational conditions,
acknowledging the non-linear correlations [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], overcoming in this way the dificulties of
physicalbased modelling. As a significant contribution, the research has embarked upon extensive data
collection from existing literature to further machine-learning eforts for predicting bio-oil
yields. This dataset exhibits the challenge of missing values, and an analysis for predicting
bio-oil yields comparing diferent imputation methods to handle missing data is provided.
      </p>
      <p>Section 2, ”Materials and Methods”, introduces the collected dataset related to waste biomasses
and bio-oils needed to enforce data-centric research and the procedures used for handling
missing data; Section 3, ”Results”, presents the comparison between diferent machine learning
algorithms for predicting the yield of bio-oil; Section 4, ”Discussion with Conclusions” presents
the conclusions elucidating the novelty of this work about the existing literature.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Material and methods</title>
      <p>The following section introduces the dataset, which has been gathered from the literature to
contribute to AI and renewable energy research in the context of waste biomasses. The missing
data in the dataset has been addressed before establishing a machine-learning benchmark for
predicting bio-oil yield, as shown in section 3. Section 2.1 presents the details of the dataset,
while Section 2.2 describes the frameworks used to handle the missing data.</p>
      <sec id="sec-3-1">
        <title>2.1. Data collection</title>
        <p>This study’s collected dataset, consisting of 1057 entries, includes proximate analysis of biomass
like ash fixed-carbon and volatiles, ultimate analysis of biomass like carbon hydrogen oxygen
and nitrogen, lignocellulosic content, and plant operative conditions as temperature, heating
rate, particle size, and nitrogen flow rate. The target variable in this research work is the bio-oil
yield.</p>
        <p>The extent of missing data varies across diferent variables. For the lignocellulosic content,
missing data account for between 40 and 50 percent of total entries, whereas for ultimate and
proximate analyses, the missing data are close to 5 percent and around 10 percent, respectively.
This prevalence of missing data necessitates the implementation of data imputation techniques
before applying models to estimate bio-oil yield.</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Missing data imputation</title>
        <p>
          In this work, the issue of missing data was addressed by leveraging four distinct imputation
techniques, utilizing an iterative method in a round-robin fashion as explained in the
scikitlearn iterative-imputer doc1, a framework inspired on [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. This procedure involves a systematic
approach to filling in the missing values for each variable iteratively, utilizing the remaining
variables as predictors. Each missing value was initially provisionally filled with rudimentary
estimates such as the variable’s mean, median, or mode. After that, one variable containing
missing values was chosen (Variable A for this instance), and its preliminarily filled values
1https://scikit-learn.org/stable/modules/impute.html#iterative-imputer
were treated as missing again. The remaining variables, now including the initially presenting
missing values but substituted with estimates in the first step, were employed to predict the
missing values for Variable A. This phase uses various imputation models: Random Forest as
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], k-nearest Neighbors (kNN) as [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], and Support Vector Regressor (SVR).
        </p>
        <p>
          Lastly, a fourth approach, the Variational Autoencoder (VAE), was used for data imputation
purposes, as seen in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. As a type of artificial neural network, the VAE generates
complex data distribution models by utilizing principles of variational inference to address the
problem of missing values.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Results</title>
      <p>Four versions of the complete dataset were created without missing values from the original
dataset using the four imputation algorithms above. Machine learning algorithms were then
employed to give a benchmark on bio-oil yields on this newly collected dataset. In Tables 1, 2,
3, and 4, the comparison of machine learning algorithms to predict the bio-oil yield is shown on
the imputed versions of the datasets using KNN, random forest, support vector regressor, and
variational autoencoder, respectively.</p>
      <p>Models PMeeracennAtabgseolEurtreor
Tree 0.1
Linear Regression 0.2
GBTree 0.1
RF 0.1
KNN7 0.2
MLP 0.2
LibSVM 0.2
RBFRegressor 0.2
KNN3 0.1
KNN5 0.1
AdditiveRegression 0.2</p>
      <p>In Tables 1, 2, 3, and 4, it is shown that the handling of missing values with the iterative
imputation using Random Forest combined with random forest for regression on bio-oil yield
produces the best results, resulting in the mean absolute error of 2.6 and adjusted R-squared
of 0.9. The models were trained using a hold-out split of 70-15-15 for training, validation, and
testing.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Discussion with Conclusions</title>
      <p>The produced bio-oil yield from biomass waste was successfully predicted with random forest
regression combined with an iterative imputation method based on Random Forest to tackle
the problem of missing data, achieving an R-squared value of approximately 0.9 and a mean
absolute error of about 2.6.</p>
      <p>
        To the best of our knowledge, this is the benchmark on bio-oil yield on the more extensive
dataset collected in literature in the context of pyrolysis of biomass, which hence contains a
broader range of biomass properties and plant operative variables, concerning other studies
on diferent and smaller datasets like [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [9] and [10]. It is worth emphasizing that in the
context of renewable energy from biomass waste, creating an extensive dataset, as was done for
the specific process of pyrolysis on which an algorithm can be trained to predict bio-oil yield
with good performance, provides a valuable tool for plant modelling, aiding in simulating new
scenarios for optimizing operative conditions and facilitating the development of systems to
handle uncertainties with their predictive capabilities, overcoming in this way the dificulties of
modelling with physics models and boosting the research to data-centric methodologies.
of bio-oil characteristics quantitatively relating to biomass compositions and pyrolysis
conditions, Fuel 312 (2022) 122812.
[9] Q. Tang, Y. Chen, H. Yang, M. Liu, H. Xiao, Z. Wu, H. Chen, S. Naqvi, Prediction of
bio-oil yield and hydrogen contents based on machine learning method: efect of biomass
compositions and pyrolysis conditions, Energy &amp; Fuels 34 (2020) 11050–11060.
[10] K. Yang, K. Wu, H. Zhang, Machine learning prediction of the yield and oxygen content of
bio-oil via biomass characteristics and pyrolysis conditions, Energy 254 (2022) 124320.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <article-title>Lignocellulosic biomass pyrolysis mechanism: A state-of-the-art review</article-title>
          ,
          <source>Progress in Energy and Combustion Science</source>
          <volume>62</volume>
          (
          <year>2017</year>
          )
          <fpage>33</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <article-title>Pyrolysis products from industrial waste biomass based on a neural network model</article-title>
          ,
          <source>Journal of Analytical and Applied Pyrolysis</source>
          <volume>120</volume>
          (
          <year>2016</year>
          )
          <fpage>94</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Van Buuren</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>Groothuis-Oudshoorn, mice: Multivariate imputation by chained equations in r</article-title>
          ,
          <source>Journal of Statistical Software</source>
          <volume>45</volume>
          (
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Stekhoven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bühlmann</surname>
          </string-name>
          ,
          <article-title>Missforest-non-parametric missing value imputation for mixed-type data</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>28</volume>
          (
          <year>2012</year>
          )
          <fpage>112</fpage>
          -
          <lpage>118</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O.</given-names>
            <surname>Troyanskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cantor</surname>
          </string-name>
          , G. Sherlock, P. Brown, T. Hastie,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Botstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Altman</surname>
          </string-name>
          ,
          <article-title>Missing value estimation methods for dna microarrays</article-title>
          ,
          <source>Bioinformatics</source>
          <volume>17</volume>
          (
          <year>2001</year>
          )
          <fpage>520</fpage>
          -
          <lpage>525</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Gevaert</surname>
          </string-name>
          ,
          <article-title>Genomic data imputation with variational auto-encoders</article-title>
          ,
          <source>GigaScience</source>
          <volume>9</volume>
          (
          <year>2020</year>
          )
          <article-title>giaa082</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Boquet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vicario</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Morell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Serrano</surname>
          </string-name>
          ,
          <article-title>Missing data in trafic estimation: A variational autoencoder imputation method</article-title>
          ,
          <source>in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>2882</fpage>
          -
          <lpage>2886</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Qian</surname>
          </string-name>
          , Machine learning prediction
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>