<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Recognition of Sports Scores using PyTessaract OCR and CNN: SportsVideo Task at MediaEval 2023</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bhuvana Jayaraman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirnalinee TT</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harshida Sujatha Palaniraj</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohith Adluru</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sanjjit Sounderrajan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Sri Sivasubramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Chennai, Tamil Nadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the dynamic realm of competitive swimming, timely and accurate dissemination of race results is paramount. The advent of digital display boards has streamlined this process, ofering real-time performance metrics for each swimmer. However, extracting and recognizing the characters displayed on these boards pose unique challenges. Our proposed methodology addresses these intricacies, ofering a robust solution for real-time score recognition amidst dynamic game-play and varying camera angles. The systematic approach leverages advanced image preprocessing techniques to refine input images, employs PyTesseract OCR for real-time interpretation of textual information from swimming board images, and integrates post-processing validation to enhance the accuracy of the extracted race results. Our findings not only contribute to the evolving field of sports analytics but also pave the way for enhanced viewer experience and enriched data accessibility in the realm of sports broadcasting.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Swimming competitions have long been a focal point of sporting events, attracting participants
and enthusiasts alike. In this era of fast-paced sporting events, the real-time availability of
race results is of paramount importance, ofering athletes insights into their performance and
spectators the thrill of immediate competition outcomes.</p>
      <p>
        Digital boards have emerged as indispensable tools in this context, acting as conduits for
promptly showcasing race results. These boards, equipped with advanced display technology,
present a visually rich representation of crucial information such as swimmer names, times,
and rankings. However, the process of manually extracting this information from the visually
diverse and dynamic context of swimming events is arduous and prone to errors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The motivation behind this research stems from the recognition of the transformative potential
of OCR technology in addressing these challenges. By automating the extraction of textual
information from digital boards, OCR not only promises to expedite the process but also ensures
a higher degree of accuracy. This, in turn, enhances the overall experience for both participants
and spectators, fostering a more eficient and error-resistant mechanism for obtaining vital race
results.</p>
      <p>
        The primary objectives of this research, within the context of Task 6 to recognize results
of races [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], involve evaluating OCR’s efectiveness in recognizing characters under diverse
conditions encountered in swimming environments and developing a tailored preprocessing
pipeline to enhance its adaptability. By achieving these objectives, the study seeks to contribute
to the eficient and accurate extraction of critical race information, ofering potential applications
beyond competitive swimming in visually complex scenarios.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Extraction and recognition of text from an image is an important step to display and predict
the results. Tesseract OCR is implemented to extract texts from given video frames or images
with an additional image processing filter, OpenCV [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Some of the processes used in are
morphological operations and Gabor filters for preprocessing, edge detection, DWT and SWT
for feature extraction and Hough transform and k-means clustering for segmentation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        For unconstrained image indexing and retrieval system using a neural network, this paper
suggests extracting a set of features from ROI for that specific color plane and using them
further in a feature-based classifier to determine if the ROI contains text or non-text blocks.
The blocks identified as text are next given as input to an OCR. The OCR output in the form of
ASCII characters forming words is stored in a database as keywords with reference for future
retrieval [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        The existing algorithms have been modified to be efective on blurred and noisy images.
Images with clear letter boundaries are filtered using OCR to detect the text [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Text extraction involves detection, localization, tracking, binarization, extraction,
enhancement and recognition of the text from the given image. The proposed methods were based on
morphological operators, wavelet transform, artificial neural network, skeletonization operation,
edge detection algorithm, histogram technique, etc [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The method used to extract text regions
for the objective of image segmentation uses DWT and k-means clustering [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Moreover,
methods such as texture-based text extraction approach to existing edge-based and connected
component (CC)-based algorithms were proposed [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for text extraction due to variables in
orientation, alignment, font, size, poor picture contrast, and complex background in the text.
      </p>
      <p>
        This paper focuses on text segmentation in textured backgrounds with similar colors to text,
outperforming recent works on challenging images.[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Approach</title>
      <p>Our comprehensive approach to the automated extraction of swimming race results integrates
image preprocessing, Optical Character Recognition (OCR) with PyTesseract, and meticulous
post-processing validation. Leveraging OpenCV and PyTesseract, we systematically optimize
input images, configure OCR settings, and validate extracted information. Simultaneously, our
innovative solution incorporates a hybrid methodology for score extraction in table tennis
and swimming events. This approach combines Convolutional Neural Network (CNN) feature
extraction with Tesseract OCR for both visual and textual cues. Initial benchmarks involved
training the CNN on visual features and parallel processing with Tesseract OCR for text
recognition. Successes include accurate score extraction and additional context, while challenges
involved varying fonts and layouts. Refined CNN architecture, preprocessing optimization,
and continuous learning mechanisms addressed these challenges. Real-world feedback and
validation on diverse datasets guided our iterative development, resulting in a robust solution
for both swimming race results and sports event scores.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset Description</title>
        <p>The dataset comprises a collection of images capturing swimming scoreboards with associated
player names and scores, sourced from various swimming competitions. The dataset includes
variations in lighting conditions, background settings, scoreboard designs, and perspectives
to enhance the development and evaluation of algorithms for accurate score extraction from
scoreboard images.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Image Preprocessing</title>
        <p>In the initial step of image preprocessing, the swimming board image is loaded using the
OpenCV library. The primary goal is to enhance the subsequent OCR process by simplifying
the image’s structure. Grayscale conversion is applied to eliminate color complexities, as the
focus is on textual information. Optional thresholding techniques are considered to create a
binary image, accentuating text while minimizing background noise. Additionally, denoising is
employed to further enhance the clarity of the text on the image. These preprocessing steps
collectively lay the groundwork for accurate OCR by optimizing the input image.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. OCR Text Extraction</title>
        <p>Following image preprocessing, the PyTesseract library is employed for Optical Character
Recognition (OCR). The OCR process involves interpreting the text present on the swimming
board image. It is crucial to configure the OCR settings appropriately, specifying the OCR
Engine Mode (OEM) and Page Segmentation Mode (PSM). These configurations are selected
based on the specific characteristics of the text in the swimming race result boards. PyTesseract
translates the visual information in the image into machine-readable text, forming the basis for
subsequent analysis.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Post Processing</title>
        <p>Post-processing steps are crucial for refining and organizing the extracted text into meaningful
race results. The extracted text is split into relevant components based on the expected structure
of the race results. Following this, custom validation checks are implemented to ensure that the
extracted information aligns with the expected format of swimming race results. Any errors or
inconsistencies introduced during the OCR process are addressed in this step. Post-processing
thus serves as a critical phase for ensuring the accuracy and reliability of the extracted race
results.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Hybrid Approach: PyTesseract OCR + CNN-Based Score Extraction</title>
        <p>Our streamlined approach starts with dataset curation, collecting diverse annotated scoreboards
from swimming events. PyTesseract OCR extracts initial text, focusing on visible scores and
details. Preprocessing involves normalizing and resizing images for consistency and model
generalization. VGG16, a pre-trained CNN, then captures visual patterns associated with scores.
A custom score layer is added that enhances the model’s ability to predict numerical scores.
Validation optimizes hyperparameters for robust performance. Real-time deployment prioritizes
CNN predictions, augmented by consolidation and confidence scoring for improved accuracy.
Continuous learning, through periodic updates, ensures adaptability to evolving scoreboards
and presentation styles. This eficient hybrid method integrates OCR with CNN for automated
score extraction in dynamic sports events.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Analysis</title>
      <p>In our research endeavor, SSN-MLRG-TEAM2 devised a hybrid PyTesseract OCR + CNN
approach, which yielded promising results for score extraction in swimming events. The overall
accuracy reached 88.05%, showcasing a notable improvement over individual PyTesseract OCR
or CNN usage.</p>
      <p>Comparative analysis revealed our hybrid approach’s unique advantages—adaptability to
diverse sports, real-time processing, and improved accuracy. Future research will explore
extensions to other sports and domains, emphasizing responsible AI practices. Overall, our
results underscore the potential of the hybrid PyTesseract OCR + CNN approach in automating
sports analytics.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Outlook</title>
      <p>Our methodology for swimming race result extraction has exhibited a commendable accuracy of
88.05% in interpreting diverse swimming board images, providing a robust foundation for
automated sports analytics. Despite successes, challenges such as variations in text layout and font
styles persist, warranting further refinement in OCR configurations and potential integration
of machine learning techniques. Looking forward, the broader applicability of our approach
extends beyond swimming competitions to other sports with similar visual data challenges.
Future research avenues include exploring real-time data streaming integration and developing a
more comprehensive sports analytics framework. Ethical considerations, particularly regarding
data privacy and bias mitigation, remain crucial focal points. As technological advancements
continue, our work serves as a stepping stone for ongoing developments in automated sports
analytics, presenting valuable insights and avenues for future exploration and improvement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Minaee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Text extraction from texture images using masked signal decomposition</article-title>
          ,
          <source>in: 2017 IEEE Global Conference on Signal and Information Processing (GlobalSIP)</source>
          , IEEE,
          <year>2017</year>
          , pp.
          <fpage>1210</fpage>
          -
          <lpage>1214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Erades</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V. B.</given-names>
            <surname>Mansencal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Péteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Morlier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dufner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Benois-Pineau</surname>
          </string-name>
          ,
          <article-title>Sportsvideo: A multimedia dataset for event and position detection in table tennis and swimming</article-title>
          ,
          <source>in: Working Notes Proceedings of the MediaEval 2023 Workshop</source>
          , Amsterdam,
          <source>The Netherlands and Online and Online, 1-2 February</source>
          <year>2024</year>
          , CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Chatterjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mondal</surname>
          </string-name>
          ,
          <article-title>Efects of diferent filters on text extractions from videos using tesseract</article-title>
          ,
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .13140/RG.2.2.36025.08804.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Manjunath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guruswamy</surname>
          </string-name>
          ,
          <source>A Survey on Text Detection from Document Images</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>961</fpage>
          -
          <lpage>972</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-15-0633-8_
          <fpage>98</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Misra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Swain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. K.</given-names>
            <surname>Mantri</surname>
          </string-name>
          ,
          <article-title>Text extraction and recognition from image using neural network</article-title>
          ,
          <source>International journal of computer applications 40</source>
          (
          <year>2012</year>
          )
          <fpage>13</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.-M.</given-names>
            <surname>Chidiac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Damien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yaacoub</surname>
          </string-name>
          ,
          <article-title>A robust algorithm for text extraction from images</article-title>
          ,
          <source>in: 2016 39th International Conference on Telecommunications and Signal Processing (TSP)</source>
          , IEEE,
          <year>2016</year>
          , pp.
          <fpage>493</fpage>
          -
          <lpage>497</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sumathi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Santhanam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Devi</surname>
          </string-name>
          ,
          <article-title>A survey on various approaches of text extraction in images</article-title>
          ,
          <source>International Journal of Computer Science and Engineering Survey</source>
          <volume>3</volume>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ghai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <article-title>A new approach to extract text from images based on dwt and kmeans clustering</article-title>
          ,
          <source>International Journal of Computational Intelligence Systems</source>
          <volume>9</volume>
          (
          <year>2016</year>
          )
          <fpage>900</fpage>
          -
          <lpage>916</lpage>
          . doi:
          <volume>10</volume>
          .1080/18756891.
          <year>2016</year>
          .
          <volume>1237189</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ghai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <article-title>Comparative analysis of multi-scale wavelet decomposition and k-means clustering based text extraction</article-title>
          ,
          <source>Wireless Personal Communications</source>
          <volume>109</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s11277-019-06574-w.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ghai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <source>Comparison of Diferent Text Extraction Techniques for Complex Color Images</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>139</fpage>
          -
          <lpage>160</lpage>
          . doi:
          <volume>10</volume>
          .1002/9781119861850.ch9.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>