<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Generation of a Book of Abstracts for Conferences that use Indico Platform</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Joint Institute for Nuclear Research</institution>
          ,
          <addr-line>Dubna</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Plekhanov Russian University of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2091</year>
      </pub-date>
      <fpage>315</fpage>
      <lpage>325</lpage>
      <abstract>
        <p>s automatically. The second requirement is quite important too. Authors may call the same a liation institutes di erently, mix di erent languages in the "authors" eld and the abstract text itself, create too short or too long abstracts. To ful ll the requirements another approach has been chosen: to use the XML representation of all abstracts, check the correctness of each of them, and create Microsoft Word document DOCX based on a template. This approach has been successfully used during the International Symposium on Nuclear Electronics and Computing 2019.</p>
      </abstract>
      <kwd-group>
        <kwd>Document generation</kwd>
        <kwd>Book of abstracts</kwd>
        <kwd>Automated system</kwd>
        <kwd>DOCX</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Automatic document generation is an important topic especially for systems that
process a big amount of data to generate standardized reports or documents.
Manual document creation may be tedious and error-prone. And if the data
amount is large and document format is not trivial the creation of a document
may become too expensive in terms of time. The worst thing is that without
some additional e orts there is no way to prove that the document is free of
typos or mistakes.</p>
      <p>The task of document generation is important for Joint Institute for Nuclear
Research(JINR) - International Intergovernmental Organization which unites
scientists from areas of nuclear physics, high-energy physics, neutron physics,
information technologies, and radiation biology. JINR organizes and hosts more
than 40 international conferences and meetings annually. Organization of a
conference requires a substantial amount of work related to conference information
publication, registration of participants, abstracts collection, time table creation
etc. For this purpose, the Indico system is widely used.</p>
      <p>
        Indico is a software started as a European project in 2002. In 2004, CERN
adopted it as its own event-management solution and has nanced its
development since. Today Indico is an Open Source Software available under MIT
license. As a tool for conference organization, Indico provides a web page with
information about the event, registration form, abstract submission form,
survey form, time schedule of the event, links to web pages with other information
related to the event, and some other features. With its rich functionality, in the
world of high energy physics, Indico became a de facto standard tool for event
organization[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>One of the features of Indico is abstract submission form and the possibility
to generate the book of abstracts automatically. All abstract texts collected as
a plain text, but with the possibility to include LaTeX formulas. The book of
abstracts is an important aspect of a scienti c conference. It should be accessible
before the talks start, so all the participants may get it, get acquainted with it
to understand which other talks may be interesting and important to them. The
book of abstracts may appear in two forms: a printed copy or a digital copy. It
is always possible to read any particular abstract on a page dedicated to a talk
if there is no book prepared in Indico.</p>
      <p>
        The task of the creation of a book of abstract is not unique to JINR. But there
are not many publications about approaches and methods that allow automating
the process. The only example has been found in publication [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Authors there
use R language and LaTeX to not only generate abstracts but also to generate
timetable.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Requirements to a Book of Abstracts</title>
      <p>Standard abstract consists of several blocks of information: title, list of authors,
list of a liations, corresponding author email, and the text of the abstract itself.
Simple example is on the Fig. 1</p>
      <p>Each author has one or several a liations. A liations marked by numbers.
The email address relates to the corresponding author. Sometimes there are
several corresponding authors. Email addresses marked by letters. The Indico
system generates rather correct abstract structure. The only issue may be that in
Indico generated abstracts the email addresses are displayed, but not connected
to a particular author.
Our primary goal was to simplify a process of a book of abstracts creation for
the International Symposium on Nuclear Electronics and Computing(NEC)
organized by JINR in 2019. Usually, a book of abstract for this conference consists
of around 150 abstracts. The creation of this book was performed manually and
required a lot of e ort. The work was monotonous, boring, and related to many
copy/paste cycles. That led to typos and errors during document creation. The
work could be divided between several people which lead to spending of around
30 man-hours in total just for one book of abstracts. Manual indexing of a
liations could lead to mistakes and required from the editor an additional check.</p>
      <p>In JINR for some conferences, the book of abstracts should be at least as
a digital copy in PDF format. And if the printed version is also required JINR
Publishing Department accepts PDF documents as a source for books. A simple
way is to generate a book using the Indico system. The result will contain the
title of the event, list of contents, and all the abstracts sorted by abstract title,
presenter name, section name, or some other features. The following
requirements make the Indico generated book of abstracts not suitable without some
additional editing:
1. After the title page, the second page with the conference annotation in
Russian and English language should follow.
2. The third page contains general conference information and topic covered
during the conference.
3. The fourth page is a list of Program Committee.
4. The fth page is a list of Organizing Committee.
5. Then the list of contents follows. But abstracts should be divided by sections
and the order of abstracts may be changed in order to bring key talks on
top of the section. Various ordering approaches may be used depending on
the conference.
6. On one page is only one abstract. Abstracts should be grouped by sections.</p>
      <p>Section titles should occupy one page, be placed at the center of the page,
and be capitalized.</p>
      <p>While the rst four requirements may be ful lled by editing a PDF document,
requirements 5 and 6 would require substantial e orts in the PDF editor and
the list of contents may be easily broken.</p>
      <p>Another important issue during the creation of a book of abstracts is the
fact that if we just put all of the abstracts from abstract submission forms in
one document the inconsistencies and errors will become visible. Among all the
di erent problems the following are more common:
{ Authors with the same a liation writes them di erently. For example with
the correct "Joint Institute for Nuclear Research" a liation, we have seen the
following variations: JINR, LIT JINR, JINR LIT, Laboratory of Information
Technologies JINR, Joint Institute of Nuclear Research, Joint Institute for
nuclear research, etc. The biggest issue here is that the a liation of the
author is asked once during the registration of the Indico account. So, the
author cannot change a liation lling the abstract submission form. To do
that author should go to the Indico pro le settings to change a liation there.
{ Abstract titles in the book of abstracts may be either fully capitalized or just
starting with the capital. Authors by themselves decide whether to write it
in capital or not. The editor of the book of abstract should check every title
manually and bring it to the right format.
{ The languages of di erent pieces of information about abstracts may be
written in a di erent language. For example, during registration, the author
wrote the name in Russian, but during abstract form submission used
English. The nal book created automatically by Indico will use Russian for
names and English for abstract text. That is a discrepancy and it should be
xed. Generally, it requires organizers to communicate with the author to
get the right spelling of the name in the language of the abstract text.
{ Abstract text is usually con ned by 250 words. Some authors may violate
this rule. The editor sometimes should manually check the number of words
and send requests to authors to shorten texts.
2.2</p>
      <sec id="sec-2-1">
        <title>Automatic Processing and Corrections of the Data for a Book of Abstracts</title>
        <p>The described requirements and issues demonstrate that the problem of the
generation of a document with some speci c format is not the only issue with
document generation. The analysis and automatic corrections of the text are
also possible and required.</p>
        <p>It is possible to correct the a liation name or capitalize on the title of an
abstract. But, mixing of languages inside an abstract usually requires organizers
to communicate with authors. To perform automatic analysis of texts for book
abstracts it is required to have a convenient source of data. Fortunately, Indico
provides the possibility to download the XML document with the representation
of all authors and abstracts. That XML document may be used to nd usual
inconsistencies. After that, the corrections may be done manually in the Indico
system. In the newer versions of the Indico, the XML format has been replaced
with JSON. In the scope of this article, we will refer only to XML since it was
the only option available for use at the moment.</p>
        <p>Once the information about all abstracts and necessary corrections are
available, it is possible to make all corrections manually and generate a document.
Or use that data to generate the nal document automatically. This would give
great exibility in terms of content and the form of the nal document.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methods of Document Generation</title>
      <p>The nal goal of document generation is the creation of a PDF le that can
be distributed as a digital copy or be printed in a printing house. The only
disadvantage of the PDF format is the fact that PDF is relatively di cult to
edit. The possibility to edit the generated document is necessary since some
manual changes may be required during the later stage of the book creation.
The generation of the PDF le directly from the data is also not as simple as a
generation of PDF from HTML, DOCX, or TEX formats. We will overview just
two common ways to generate PDF: from TEX le using the LaTeX document
preparation system, and from DOCX le using Microsoft Word word processor.
3.1</p>
      <sec id="sec-3-1">
        <title>Using the LaTeX System</title>
        <p>LaTeX is the standard for the publication of scienti c documents. It provides
high-quality document printing, so the document looks like a book. The system
allows generating a document with speci ed formatting and the possibility to
draw mathematical formulas. In our case, it is possible to draw mathematical
formulas directly from abstracts texts.</p>
        <p>The biggest advantage of this method is the possibility to generate a TEX
le without third-party libraries. Once the template TEX le is available it is
rather easy to ll it with text form source XML le. And LaTeX system itself is
free software and could be used to generate documents without any payments.
And some modern TEX editors even support WYSIWIG mode (What You See
Is What You Get, means that editing software allows content to be edited in
a form that resembles its appearance when printed or displayed as a nished
product), although this type of editors may be non-free.</p>
        <p>However, the preparation of a template in the LaTeX system usually requires
some special skills. Another issue with LaTeX was the fact that not all organizing
committees had a LaTeX template available. Some committees used a complex
DOCX template with macros to apply correct formatting to di erent parts of a
text.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Using a Prepared DOCX Template</title>
        <p>This method involves the use of a prepared DOCX template. Templates may
be di erent. It is possible to make it simple and just de ne several types of
formatting for di erent elds, like "Abstract title", "Authors", "Abstract text".
Sometimes more complex templates with macros may be used. The resulting
document may be generated from a template using special third-party libraries.
The generated DOCX document may be additionally edited in Microsoft Word.
The WYSIWIG is originally supported by Microsoft Word and may simplify the
manual editing process.</p>
        <p>The disadvantage of this approach is the need to use third-party libraries.
There is no guarantee that they will always generate the correct document and
the newer version of Microsoft Word may render generated DOCX les di
erently. Microsoft Word itself is non-free software, however, currently, it is a
standard software for document creation in many organizations, including JINR. So,
everybody has access to Microsoft Word. Another issue is the support of LaTeX
formulas. There are some approaches to include them in DOCX documents but
most of them require some manual operations.</p>
        <p>There were three reasons for us to use DOCX templates for the generation
of the book of abstracts. First, the DOCX template has already been created
and used for several previous NEC conferences. Second, libraries for Python
language for work with DOCX templates and documents have been found, and
their functionality was proved during initial tests. Third, in our case, members
of the organizing committees responsible for the books of abstracts preferred
Microsoft O ce.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Implementation of the Book of Abstract Generation</title>
    </sec>
    <sec id="sec-5">
      <title>System</title>
      <p>4.1</p>
      <sec id="sec-5-1">
        <title>Tools and Technologies Used for Implementation</title>
        <p>The Python3 language has been chosen as a primary language for the developed
system. It is possible to use programs developed in Python under Linux and
Windows operating systems. Moreover, Python provides the possibility to make
the developed program available as a web-application. The Python is quite
popular in science organizations and the developed system may be easily used and
changed by other users.</p>
        <p>
          The following Python libraries were used in the developed system:
{ xml.etree.cElementTree is a Python standard library written in C for
extracting data from XML les. Now, since Indico system introduced export
in JSON format the json library from standard Python libraries may be
used.
{ python-docx is Python library for reading, writing and creating sub
documents les [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
{ python-docx-template is a library for modifying DOCX documents that have
already been created. python-docx-template was created because python-docx
is powerful for creating documents, but not for modifying them. This package
itself based on python-docx [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>The developed system uses the Command Line Interface. It has been chosen
since it is much easier to make it look the same on both Linux and Windows
operating systems. The possible next step is providing a developed system as a
service through a Web Interface.
4.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>General Scheme of the Program</title>
        <p>The generation of a book of abstracts is done in several steps:
1. Export of the XML le with all abstracts data.
2. Parsing of XML le in Python and creating object with all data from XML
le.
3. Automatic correction of standard inconsistencies described in section 2.1.
4. Displaying the noti cations about issues that cannot be xed automatically.
5. Generation of the "preface" part with general conference information
described in list of requirements items 1-4 section 2.1.
6. Generation of abstracts part which contain only abstracts.
7. Creation of the nal document by concatenating preface and abstract parts.</p>
        <p>To allow the execution of these steps several additional les are required:
templates, information about the conference to be used for preface generation,
CSV le with validated a liation names.</p>
        <p>Preparation of a DOCX Template for Preface. The preface template will
determine the look and layout of the preface part of the book. For that, an
existing book of abstracts from a previous conference is used. Instead of the
information which changes between conferences, we insert variables enclosed in
curly brackets. Each variable must have a unique name. If a variable with the
same name is repeated in several places, the text associated with it will be
inserted instead of all of them. In places where the information may change (like
endings of ordinal numbers in conference title) and can not be automatically
inserted, the Microsoft O ce Word notes used. They show places where the
editor's attention is required. Styles should be chosen and applied during
template preparation because they will be used in the nal document of the book
of abstracts. The example of a preface template is shown in Fig. 2.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Making an Additional XML File with Data for Preface part. To ll the</title>
        <p>previously described template additional XML le should be created. It contains
general information about the conference (conference name and number, date,
description) in two languages (see Fig. 3). It should be done once for every book
of abstracts.</p>
        <p>Fig. 2. Example of a DOCX template for preface.
Making a CSV File. Since the same a liations may be written by di erent
authors di erently it is important to bring them to consistency at least in one
particular book of abstracts. For that purpose, the CSV le is used. It works
like a dictionary. All ways to write a particular a liation should be included in
this le. If the name of a liation used inside an abstract description does not
exist in the CSV le, the warning message appears and suggests to add it to the
le. If the spelling of the a liation is correct it still has to be in the CSV le
and refer to itself. This ensures that the editor con rms the correction of the
a liation name. An example of a text from the described CSV le is shown in
Fig. 4.</p>
        <p>When parsing an input XML le with abstracts, the program accesses CSV
le and compares the spelling of the current A liation with the correct one. If
it matches one of the incorrect variants, the program replaces it with the correct
one and inserts it to the nal document in an appropriate place. If no such
A liation is found in the CSV le, the program displays a message stating that
this A liation is not in the CSV standards le.</p>
      </sec>
      <sec id="sec-5-4">
        <title>Checking the Language Consistency. To ensure that di erent parts of the</title>
        <p>same abstract are not written in di erent languages, we created a special function
checking and indicating the language inconsistency (see Fig. 5).
Checking the Amount of Words. The program checks the number of words
in the abstract. If it is greater or lesser than the speci ed limit, the warning is
displayed in the log of the execution.
4.3</p>
      </sec>
      <sec id="sec-5-5">
        <title>Making corrections</title>
        <p>Since many issues determined by the developed system can not be resolved
automatically the editor should resolve them manually. There are at least three ways
to perform corrections: in the Indico system itself, in some con guration les
(now we have two: XML with conference info and CSV with correct a liations),
in the generated document.</p>
        <p>Fixing the data in the Indico system is the best way. Usually, it requires some
work from authors of the abstract or at least their agreement. All the changes will
be included in the generated book of abstracts during the next generations. This
works well with inconsistent language use or incorrect abstract text size. If the x
inside Indico is impractical, like with a liation names, the special con guration
le may be introduced to perform corrections during every generation of the
particular book of abstracts. This is what has been done with a liation names.</p>
        <p>The xing data inside the generated document is a viable option, but only
when everything else is established and there are no new generations expected.
That is because all changes inside could not be used in the next generation and
should be done manually again. In real use-case during the creation of a book
of abstracts for NEC conference all of the described approaches have been used.
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <p>The work involved designing and creating an automated abstract book generator
that would be identical to the one created by hand. Besides, we had the task to
check the source XML le for possible errors listed in this article. As a result, a
software product was created that receives four necessary input les:
{ an XML le exported from Indico containing abstracts,
{ an XML le containing conference information,
{ a CSV le containing the A liations writing standards,
{ a DOCX template that will be used as a base to generate the nal document.</p>
      <p>As a result, the program creates a nal DOCX le with the formatting used in
the template. Besides, the program performs checks for possible errors described
in this article. Using the information obtained during checks, an editor of the
book of abstracts can decide how to correct mistakes: make changes in Indico,
ask for a new feature for the book of abstracts generation tool, or change the
nal document directly.</p>
      <p>The applied approach greatly reduced the amount of e ort required to
generate a document of a book of abstracts. The pursuit to simplify manual indexing
and copying led to introducing automatic checks and corrections. And
developed system allowed to generate a document in a format that is known to the
end-user.</p>
      <p>
        The developed product is a command-line interface (CLI) application written
in Python v3.5 for Linux and Windows. This requires some speci c libraries to
be installed on the client machine before the use of the system. All source code
is available under GNU General Public License v3.0 at [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The software product, methods, and approaches described in this article were
used during the preparation of the book of abstracts for the NEC'2019
Symposium. The generated book is available at [6].</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Indico</given-names>
            <surname>Project</surname>
          </string-name>
          , https://docs.getindico.io/en/latest/.
          <source>Last accessed 25 July 2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Lara</given-names>
            <surname>Lusa</surname>
          </string-name>
          , Andrej Blejec,
          <source>Automated Preparation of the Book of Abstracts for Scienti c Conferences using R and LaTeX: Infor Med Slov</source>
          <year>2009</year>
          ,
          <volume>14</volume>
          (
          <issue>1-2</issue>
          ), pp.
          <volume>10</volume>
          {
          <fpage>18</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>3. python-docx library documentation</article-title>
          , https://pythondocx.readthedocs.io/en/latest/.
          <source>Last accessed 25 July 2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>4. python-docx-template library documentation</article-title>
          , https://docxtpl.readthedocs.io/en/latest/.
          <source>Last accessed 25 July 2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Project GitHub repository https://github.com/trnkv/IndicoAbstract. Last accessed
          <issue>25</issue>
          <year>July 2020</year>
          6.
          <article-title>Generated book of abstracts for NEC2019 conference</article-title>
          https://indico.jinr.ru/event/738/attachments/4884/6443/NEC 2019 BoA.pdf accessed
          <issue>25 July 2020</issue>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>