<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Automated Text Summarization in Clinical Approaches Trials: Towards Explainable AI Solutions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zakaria Benzadri</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maroua Benhamlaoui</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Naila Marir</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Efat College of Engineering, Efat University</institution>
          ,
          <addr-line>Jeddah</addr-line>
          ,
          <country country="SA">Saudi Arabia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University Fréres Mentouri- Constantine 1</institution>
          ,
          <addr-line>Constantine</addr-line>
          ,
          <country country="DZ">Algeria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Constantine 2 - Abdelhamid Mehri, LIRE Laboratory</institution>
          ,
          <addr-line>Constantine</addr-line>
          ,
          <country country="DZ">Algeria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The exponential rise in digital content has made Automatic Text Summarization (ATS) an indispensable tool in information retrieval and knowledge management. Neural network architectures have significantly advanced the eficiency and accuracy of ATS systems, particularly in specialized domains requiring precise and reliable information synthesis. This paper provides a comprehensive review of recent advancements in neural network-based summarization models, highlighting the state-of-the-art methodologies, challenges faced in adapting these models to complex, domain-specific texts, and areas for future research. Additionally, the limitations of current techniques are critically analyzed, alongside the emerging need for more robust evaluation metrics to ensure the practical relevance and efectiveness of ATS systems. This work aims to provide valuable insights to researchers and practitioners dedicated to advancing the capabilities of ATS technologies.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Automatic Text Summarization</kwd>
        <kwd>Neural Networks</kwd>
        <kwd>Clinical Text Summarization</kwd>
        <kwd>Explainable AI</kwd>
        <kwd>Transparent AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The explosive growth of digital data presents significant challenges in managing, retrieving, and
synthesizing relevant information eficiently. In response, ATS has emerged as a indispensable
tool for distilling large volumes of text into concise, meaningful summaries. By providing
users with simplified representations of complex content, ATS systems can greatly enhance
information accessibility and utility. However, the complexity and domain-specific nature of
certain texts, especially in technical fields, create a demand for more sophisticated summarization
techniques.</p>
      <p>Recent advancements in neural network architectures, particularly deep learning models such
as recurrent neural networks, transformers and generative AI, have revolutionized the field of
ATS. These models, which leverage techniques such as attention mechanisms and pre-training
on large corpora, have demonstrated substantial improvements over traditional rule-based or
statistical methods. By enabling more context-aware and coherent summaries, neural networks
have set new standards for both extractive and abstractive summarization tasks.</p>
      <p>Despite these advancements, numerous challenges remain. Neural network models, while
powerful, often struggle with domain-specific content where terminological precision and
factual consistency are critical. This is particularly evident in fields where summarization tasks
must meet high standards of accuracy and reliability. Moreover, the inherent complexity of
neural networks makes them prone to issues like overfitting and generalization errors, especially
when applied to specialized corpora. These limitations call for further research into more robust,
adaptable architectures that can handle the intricacies of complex textual data.</p>
      <p>As the field continues to evolve, there is also a growing need for more comprehensive
evaluation metrics that go beyond surface-level text matching. While metrics like ROUGE and
BLEU have been widely adopted, they are often inadequate for assessing the deeper semantic
accuracy required in domain-specific summarization. Consequently, developing new metrics
that better capture the relevance, accuracy, and completeness of summaries will be crucial for
advancing the efectiveness of ATS systems in specialized applications.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Review: Neural Network Models for Clinical Text</title>
    </sec>
    <sec id="sec-3">
      <title>Summarization</title>
      <p>
        In recent years, the volume of clinical data has increased exponentially, creating a need for
eficient methods to distill key information from complex medical texts [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Clinical text
summarization, which involves condensing lengthy medical documents such as patient records,
discharge summaries, and radiology reports into concise, meaningful summaries, is essential for
improving healthcare delivery. This section presents an overview of the advancements in neural
network models for clinical text summarization, highlighting their applications in the medical
ifeld, challenges encountered, and current limitations. The focus will be on understanding
the specific characteristics of neural network-based approaches, such as sequence-to-sequence
models, transformer-based architectures, and their eficacy in extracting meaningful insights
from clinical texts. Additionally, the section will address the need for explainable AI techniques
to enhance the transparency and trustworthiness of these models in the healthcare context.
      </p>
      <sec id="sec-3-1">
        <title>2.1. Key Neural Network in Clinical Text Summarization</title>
        <p>
          Recent advancements in neural networks have significantly impacted the field of text
summarization. Sequence-to-sequence (Seq2Seq) models, transformers, BERT, and GPT are among the
most prominent models used for this purpose. Seq2Seq models, particularly those enhanced
with attention mechanisms, have shown substantial improvements in generating coherent
summaries [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Transformers, with their self-attention mechanisms, have revolutionized text
summarization by enabling the processing of entire sequences simultaneously, leading to more
contextually aware summaries [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. BERT, a bidirectional transformer, has been particularly
efective in capturing contextual information, making it a powerful tool for both extractive and
abstractive summarization tasks [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ]. GPT, with its autoregressive nature, excels in generating
human-like text, making it suitable for abstractive summarization [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          In the medical domain, these models have been adapted to handle the unique challenges posed
by clinical texts. For instance, BERT-based models have been employed to classify and summarize
clinical reports, achieving high accuracy in identifying key findings from radiographic reports
[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The Biomed-Summarizer framework integrates deep neural networks to provide
contextaware summarization of biomedical texts, ensuring that the summaries are both precise and
clinically relevant [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Additionally, models like T-BERTSum have been developed to handle
long text dependencies and latent topic mapping, which are crucial for summarizing extensive
clinical documents [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Challenges in Neural Network based Models for Clinical Text</title>
      </sec>
      <sec id="sec-3-3">
        <title>Summarization</title>
        <p>
          Clinical text summarization faces several unique challenges, including the complexity of medical
jargon, the presence of regulatory language, and the sensitivity of information. Medical texts
often contain specialized terminology that can be dificult for general models to understand and
accurately summarize [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Furthermore, the need to comply with regulatory standards adds an
additional layer of complexity, as the summaries must be both accurate and legally compliant.
        </p>
        <p>
          Current models, while advanced, still face limitations such as generalization and overfitting.
For example, BERT’s performance can be inhibited by its pretraining and tokenization methods,
which may not be well-suited for the specific nuances of clinical texts [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Additionally, models
like GPT-3.5 and ChatGPT, despite their capabilities, can generate factually inconsistent
summaries and struggle with identifying salient information in longer texts [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. These limitations
highlight the need for further refinement and adaptation of these models to better handle the
intricacies of clinical text summarization.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Explainable AI (XAI) in Clinical Decision Support</title>
      <p>Explainable AI (XAI) refers to the development of machine learning models that provide
transparent, understandable, and interpretable explanations for their decisions and predictions.
In clinical decision support, XAI plays a critical role by ofering insights into the reasoning behind
AI-driven recommendations, making them more accessible and trustworthy for healthcare
professionals. The ability to interpret AI models’ outputs is particularly important in healthcare,
where decisions can directly impact patient outcomes. This section explores the application of
XAI in clinical decision support systems (CDSS) and its potential to enhance decision-making
processes in healthcare settings.</p>
      <sec id="sec-4-1">
        <title>3.1. Applications of XAI in Clinical Decision Support</title>
        <p>
          XAI techniques are increasingly being applied to clinical decision support systems to enhance
the interpretability of AI models [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ]. These systems assist clinicians in making decisions
related to diagnosis, treatment recommendations, and patient risk assessments[
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ]. One of
the key areas where XAI is applied is in risk prediction and personalized medicine. AI models
are often used to predict patient outcomes, such as the likelihood of disease progression or
complications. XAI methods, such as SHAP (SHapley Additive exPlanations) and LIME (Local
Interpretable Model-agnostic Explanations), can identify which factors, such as test results,
demographics, or medical history, contributed most to the model’s prediction. This information
helps clinicians better understand the basis for risk assessments and make more personalized
treatment decisions.
        </p>
        <p>
          XAI is also crucial for diagnostic decision support[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. By providing reasoning behind an AI
model’s predictions, XAI helps clinicians understand how a model arrived at a specific diagnosis.
For instance, if an AI model suggests a diagnosis of pneumonia, it can explain that the prediction
was based on factors such as chest X-ray images, temperature readings, and cough symptoms.
This transparency allows healthcare professionals to validate the AI’s recommendations and
ensure the accuracy of their clinical decisions.
        </p>
        <p>
          Another area where XAI plays a critical role is in clinical text summarization[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. In clinical
environments, summarizing medical records, radiology reports, or discharge summaries can be
a time-consuming task. Transformer-based models, like BERT or GPT, are frequently used to
generate summaries, and XAI techniques, such as attention mechanisms, can highlight which
parts of the text, such as key phrases or sentences, most influenced the summary. This improves
the transparency of AI-driven text processing in clinical contexts, making the results easier to
interpret and trust.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Overview of Existing XAI Techniques</title>
        <p>
          Several XAI techniques, such as SHAP (SHapley Additive exPlanations) and LIME (Local
Interpretable Model-agnostic Explanations), have been developed to provide transparency in
clinical decision-making and enhance the interpretability of AI models. SHAP and LIME are
model-agnostic approaches that explain individual predictions by quantifying the contribution
of each feature. In healthcare, these techniques can be used to explain risk scores or diagnosis
predictions based on patient data. For example, SHAP can show how a patient’s age, blood
pressure, and previous medical history contribute to a prediction of heart disease risk, while
LIME can generate interpretable explanations for specific predictions by approximating the
model locally. Attention mechanisms, often used in deep learning models, highlight which
parts of the input data (e.g., text or images) are most important for a prediction. In clinical
settings, attention mechanisms can be applied to models processing medical texts or images to
pinpoint key features influencing decisions, providing clinicians with a visual representation of
the model’s focus areas. Additionally, counterfactual explanations ofer insights by showing
how changes in input data would lead to diferent model outcomes. In clinical decision support,
counterfactuals can demonstrate how modifying a patient’s treatment plan or clinical data would
afect the risk prediction or diagnostic outcome, helping clinicians understand the potential
impact of various decisions. These techniques, applied to NLP models, are crucial in enhancing
the transparency of AI-driven decision-making in healthcare, fostering trust and better decision
support for clinicians [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Gaps in Explainability for Neural Networks in Text Summarization</title>
        <p>Despite the significant progress in the development of neural network models for clinical text
summarization, there remain notable gaps in explainability that hinder the full adoption and
trust of these models in clinical settings. One of the primary challenges is the "black-box" nature
of many advanced neural network models, which makes it dificult for clinicians to understand
how the model generates summaries or selects key information. This lack of transparency can
undermine the reliability of AI systems, especially in critical clinical contexts where errors can
have serious consequences.</p>
        <p>A key gap is the limited interpretability of attention mechanisms, which are widely used in
models like transformers for summarization tasks. While attention mechanisms can identify
which parts of the text are most influential in generating a summary, they do not provide a
clear rationale for why certain sections were prioritized over others. For example, a model may
highlight specific symptoms or treatments in a summary, but without a deeper explanation, it
remains unclear whether the model is making decisions based on medical relevance, statistical
patterns, or other factors. This opacity can lead to confusion and reduce clinician confidence in
the system’s recommendations.</p>
        <p>Additionally, current techniques for model interpretation often struggle to account for the
complex and context-dependent nature of medical texts. Clinical documents, such as patient
histories and discharge summaries, contain nuanced information that can be influenced by a
variety of factors, including the patient’s unique medical history, current health status, and
environmental factors. Neural networks may aggregate these data points in ways that are not
always intuitive, and existing interpretability methods may fail to capture the full context of
the decision-making process.</p>
        <p>Furthermore, while models such as BERT and GPT have demonstrated strong performance in
generating clinical summaries, the explanations of these models often do not align with human
reasoning. Clinicians may struggle to relate the model’s output to their own diagnostic process,
which is typically based on years of training and clinical experience. The mismatch between
AI-generated summaries and clinical intuition can create barriers to efective collaboration
between AI systems and healthcare professionals.</p>
        <p>Finally, the absence of standardized frameworks for evaluating the explainability of clinical
text summarization models contributes to the gap in understanding. Without clear metrics
or guidelines for assessing the transparency of these models, it is dificult for researchers and
clinicians to gauge the efectiveness of XAI techniques or to compare diferent models across
studies.</p>
        <p>Addressing these gaps is crucial for improving the trustworthiness and efectiveness of neural
network-based clinical text summarization systems. By developing more interpretable models,
enhancing the transparency of decision-making processes, and aligning AI outputs with human
reasoning, researchers can create more reliable and user-friendly tools for clinicians.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Evaluation and Metrics for Clinical Text Summarization</title>
      <p>
        The efectiveness of neural network models in automatic text summarization is determined not
only by their ability to generate coherent and concise summaries but also by how well they
capture essential information from the source text. In the clinical domain, the accuracy and
relevance of the generated summaries are paramount, given the complexity and sensitivity of
medical data [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Therefore, appropriate evaluation metrics are necessary to ensure that these
models meet the high standards required for clinical applications.
      </p>
      <p>
        In this section, we explore various methods for evaluating summarization models, including
widely adopted metrics in general text summarization, as well as specialized metrics tailored
to the clinical domain. Additionally, with the growing emphasis on explainable AI (XAI), we
discuss metrics that assess the interpretability and transparency of model outputs, which are
critical for the practical use of AI systems in healthcare [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. This comprehensive evaluation
framework is crucial for identifying the strengths and limitations of existing models and guiding
future improvements in neural network architectures for clinical text summarization.
      </p>
      <sec id="sec-5-1">
        <title>4.1. Evaluating Neural Networks in Text Summarization</title>
        <p>The evaluation of neural networks in automatic text summarization is a crucial component in
determining the efectiveness of models across diferent tasks. Both qualitative and quantitative
metrics are used to assess how well a model captures the salient points of the source text and
generates coherent, accurate summaries. In the context of clinical text summarization, where
precision and information retention are paramount, robust evaluation metrics are essential for
validating the performance of these models.</p>
        <sec id="sec-5-1-1">
          <title>4.1.1. Common Evaluation Metrics</title>
          <p>
            Several standard metrics are employed to evaluate text summarization models, most notably:
• ROUGE (Recall-Oriented Understudy for Gisting Evaluation): ROUGE is one of
the most widely used metrics in text summarization, especially for extractive methods. It
measures the overlap between the generated summary and reference summaries, primarily
focusing on n-grams, word sequences, and word pairs [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. ROUGE-n, ROUGE-L (longest
common subsequence), and ROUGE-W (weighted longest common subsequence) variants
are often used. While useful, ROUGE can sometimes overestimate similarity due to
surface-level matching, making it less efective in evaluating abstractive summaries [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ].
• BLEU (Bilingual Evaluation Understudy): Originally designed for machine translation,
BLEU scores are also employed in text summarization. BLEU evaluates how closely a
generated text matches a reference text based on the precision of n-grams [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ]. However,
BLEU is less commonly used in summarization compared to ROUGE, as it often favors
shorter and more concise summaries, which might not be ideal for tasks requiring detailed
content preservation.
• METEOR (Metric for Evaluation of Translation with Explicit ORdering): METEOR
is another metric that, like BLEU, is used primarily in machine translation but has
applications in text summarization. METEOR accounts for synonymy and stemming, making it
more flexible than ROUGE and BLEU for abstractive summaries. It assigns higher scores
to summaries that capture meaning, even when surface-level similarity is lower [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ].
          </p>
        </sec>
        <sec id="sec-5-1-2">
          <title>4.1.2. Domain-Specific Evaluation in Clinical Summarization</title>
          <p>
            In clinical summarization, accuracy and relevance of the information retained in summaries are
critical. Therefore, in addition to the aforementioned metrics, several domain-specific evaluation
methods are utilized:
• Precision and Recall: These metrics are particularly important in clinical
summarization, where false positives (irrelevant information) and false negatives (missing key
clinical details) can have serious consequences. Precision evaluates the relevance of the
information included, while recall assesses the completeness of the summary [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. High
precision and recall are necessary to ensure the generated summaries provide accurate
and complete overviews of clinical reports.
• F1-Score: As a harmonic mean of precision and recall, the F1-score provides a balanced
metric to evaluate models in clinical domains where both over-inclusion of irrelevant
details and omission of critical information are unacceptable. This metric is particularly
useful when optimizing both precision and recall [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ].
• Clinical Relevance and Expert Review: Clinical text summarization models are often
evaluated through expert review, where medical professionals assess the relevance and
quality of the generated summaries [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. These evaluations are crucial, as automated
metrics alone may not fully capture the clinical importance of the information presented.
For example, a summary that scores well on ROUGE may still fail to meet clinical relevance
if key findings are omitted.
          </p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Explainability and Interpretability Metrics</title>
        <p>
          With the increasing demand for explainable AI (XAI) in healthcare, new metrics are being
developed to evaluate the transparency and interpretability of neural network models. Some
emerging metrics include:
• SHAP (SHapley Additive exPlanations): SHAP values provide insight into the
contributions of individual features or words in a model’s decision-making process, allowing
healthcare professionals to understand why certain text segments were included or
omitted in the summary [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
• LIME (Local Interpretable Model-agnostic Explanations): LIME generates local
approximations of a model’s behavior, explaining why specific summary decisions were
made. This is particularly useful in clinical settings, where trust in AI systems depends
on the interpretability of the model’s output [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
• Transparency and Trustworthiness Scores: In clinical settings, AI models are
increasingly evaluated for their interpretability using trustworthiness scores based on qualitative
assessments by medical professionals. These scores help quantify the level of trust that
clinicians can place in the AI’s decisions, which is vital for the adoption of these models
in practice [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>4.3. Challenges in Evaluation</title>
        <p>While existing metrics provide valuable insights into model performance, there are still
challenges in evaluating text summarization models, particularly in the clinical domain. For instance,
most automated metrics like ROUGE and BLEU do not adequately capture semantic accuracy
or the clinical importance of information. Moreover, the subjective nature of clinical relevance
often requires expert validation, which can be resource-intensive. Additionally, the increasing
complexity of neural network models makes explainability a growing concern, as traditional
metrics do not account for the interpretability of model decisions. Therefore, future evaluation
eforts should focus on developing more comprehensive metrics that combine automated
evaluation with expert review, while also incorporating transparency and interpretability measures
to ensure the clinical viability of summarization systems.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion: Bridging the Gap Between AI Technology and</title>
    </sec>
    <sec id="sec-7">
      <title>Clinical Practice</title>
      <sec id="sec-7-1">
        <title>5.1. Key Findings</title>
        <p>
          The studies reviewed demonstrate significant advances in the automatic summarization of
clinical trial data, both for extractive and abstractive methods. The TextRank algorithm
performed well for extractive summarization tasks, ofering concise, meaning-preserving synopses
of clinical trial descriptions. Similarly, ontological and domain-specific models, such as those
incorporating BERTSUMEXT and other neural models, have demonstrated improvements in
abstractive summarization, as seen in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
        <p>A key advantage in this domain is the use of neural architectures, particularly transformers
and pre-trained models, to address the challenges of clinical text summarization. In particular,it
achieved strong results with multi-objective optimization techniques, leveraging sentence
position and similarity metrics like: Term Frequency-Inverse Document Frequency (TF-IDF )
and Word Mover’s Distance (WMD). These methods not only outperform human gold standard
summaries in some cases but also ofer practical utility in new disease settings, such as during
the emergence of novel diseases like COVID-19.</p>
        <p>One common challenge noted in many papers was the tendency of abstractive models to
introduce factual inaccuracies, particularly when summarizing biomedical evidence across
multiple documents.</p>
      </sec>
      <sec id="sec-7-2">
        <title>5.2. Limitations</title>
        <p>While the advancements in automatic summarization are promising, there remain significant
limitations. First, the lack of transparency and controllability in many neural models limits their
practical utility in medical settings. Although systems like TrialsSummarizer aim to enhance
user verification by allowing traceability of generated tokens, more work is needed to ensure
end-users can reliably trust the outputs.</p>
        <p>Another critical limitation lies in the evaluation methods for summarization quality. The
heavy reliance on ROUGE scores is not always aligned with the preferences of healthcare
professionals. This gap between automatic and manual evaluations highlights the need for more
domain-specific evaluation metrics that better reflect clinical relevance and accuracy. Moreover,
inter-rater agreement among human evaluators often varies.</p>
      </sec>
      <sec id="sec-7-3">
        <title>5.3. Future Directions</title>
        <p>
          Looking ahead, the integration of Explainable AI (XAI) approaches holds great potential for
improving the transparency and usability of clinical summarization models. For instance,
enhancing the traceability of neural model outputs—where users can see which portions of the
input data contribute to each generated sentence—could improve trust in automated summaries.
Additionally, further fine-tuning models with clinical-specific ontologies (e.g., UMLS) could
improve both the factual accuracy and content relevance of generated summaries, as demonstrated
in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
        <p>Another promising direction is the exploration of multi-document summarization systems
that can handle vast datasets, such as those found on ClinicalTrials.gov. For instance, the EXACT
system shows that automated extraction can drastically reduce the time needed to summarize
and analyze clinical trial data for meta-analyses.</p>
        <p>Lastly, improving factual verification mechanisms within neural models should be prioritized,
as this will be crucial for the real-world deployment of such systems in clinical decision-making
contexts. Collaborating with medical professionals to ensure practical utility and relevance will
be key in advancing from experimental models to operational clinical tools.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>6. Conclusion</title>
      <p>This paper has provided an in-depth exploration of the role of ATS in clinical applications,
particularly within the context of neural network-based models. We reviewed the key
advancements in these models, emphasizing their potential to enhance clinical decision-making
processes by improving the eficiency and accuracy of summarizing complex medical texts.
Despite the significant progress made, several challenges remain, including the adaptation of
neural networks to the specific demands of the healthcare domain and the need for greater
transparency in the decision-making process.</p>
      <p>The paper also highlighted the growing importance of Explainable AI (XAI) in clinical
decision support. By incorporating XAI techniques, such as SHAP, LIME, and attention mechanisms,
into clinical text summarization, AI systems can provide valuable insights into their
predictions, fostering trust and aiding clinicians in making informed decisions. However, the gap
in explainability for neural networks, particularly in clinical text summarization, persists and
demands further investigation. Furthermore, we discussed the current limitations of existing
evaluation metrics and the need for domain-specific measures that better capture the nuances
of clinical text summarization. As medical data becomes more complex, the development of
robust evaluation frameworks will be critical to ensuring that ATS systems deliver high-quality,
clinically relevant summaries.</p>
      <p>In conclusion, while there is great potential for neural network-based ATS systems in clinical
settings, future research must address the challenges related to model explainability, evaluation
metrics, and the practical application of these technologies. By doing so, we can further enhance
the efectiveness of ATS systems and support clinicians in delivering optimal patient care.</p>
    </sec>
    <sec id="sec-9">
      <title>7. Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used the free version of ChatGPT for grammar
and spelling checks, as well as for some reformulations to improve readability.</p>
      <p>After using this tool, the authors reviewed and edited the content as needed and take full
responsibility for the publication’s content, in accordance with the for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdulnazar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Roller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schulz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kreuzthaler</surname>
          </string-name>
          ,
          <article-title>Large language models for clinical text cleansing enhance medical concept normalization</article-title>
          ,
          <source>IEEE Access 12</source>
          (
          <year>2024</year>
          )
          <fpage>147981</fpage>
          -
          <lpage>147990</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2024</year>
          .
          <volume>3472500</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Baykara</surname>
          </string-name>
          , T. Güngör,
          <article-title>Turkish abstractive text summarization using pretrained sequenceto-sequence models</article-title>
          ,
          <source>Natural Language Engineering</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K. T.</given-names>
            <surname>Chitty-Venkata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Emani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vishwanath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Somani</surname>
          </string-name>
          ,
          <article-title>Neural architecture search for transformers: A survey</article-title>
          , IEEE Access (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Al-Nabhan</surname>
          </string-name>
          ,
          <article-title>T-bertsum: Topic-aware text summarization based on bert</article-title>
          ,
          <source>IEEE Transactions on Computational Social Systems</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Moradi</surname>
          </string-name>
          , G. Dorfner,
          <string-name>
            <given-names>M.</given-names>
            <surname>Samwald</surname>
          </string-name>
          ,
          <article-title>Deep contextualized embeddings for quantifying the informative content in biomedical text summarization, Computer Methods</article-title>
          and Programs in Biomedicine (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Idnay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nestor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Soroush</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Elias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ding</surname>
          </string-name>
          , G. Durrett,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Rousseau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <article-title>Evaluating large language models on medical evidence summarization</article-title>
          ,
          <source>NPJ Digital Medicine</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Bressem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Adams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gaudin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tröltzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hamm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Makowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schüle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vahldiek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Niehues</surname>
          </string-name>
          ,
          <article-title>Highly accurate classification of chest radiographic reports using a deep learning natural language model pre-trained on 3.8 million text reports</article-title>
          ,
          <source>Bioinformatics</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Afzal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Malik</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. Malik,</surname>
          </string-name>
          <article-title>Clinical context-aware biomedical text summarization using deep neural network: Model development and validation</article-title>
          ,
          <source>Journal of Medical Internet Research</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alawad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Gounley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schaeferkoetter</surname>
          </string-name>
          , H.
          <article-title>-</article-title>
          <string-name>
            <surname>J. Yoon</surname>
          </string-name>
          , X.
          <string-name>
            <surname>-C. Wu</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Durbin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Doherty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Stroup</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Coyle</surname>
          </string-name>
          , G. Tourassi,
          <article-title>Limitations of transformers on clinical text classification</article-title>
          ,
          <source>IEEE Journal of Biomedical and Health Informatics</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N. Y.</given-names>
            <surname>Murad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Azam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Yousuf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Yalli</surname>
          </string-name>
          ,
          <article-title>Unraveling the black box: A review of explainable deep learning healthcare techniques</article-title>
          ,
          <source>IEEE Access 12</source>
          (
          <year>2024</year>
          )
          <fpage>66556</fpage>
          -
          <lpage>66568</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2024</year>
          .
          <volume>3398203</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nauman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Almadhor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Akhtar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alghuried</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alhudhaif</surname>
          </string-name>
          ,
          <article-title>Guaranteeing correctness in black-box machine learning: A fusion of explainable ai and formal methods for healthcare decision-making</article-title>
          ,
          <source>IEEE Access 12</source>
          (
          <year>2024</year>
          )
          <fpage>90299</fpage>
          -
          <lpage>90316</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2024</year>
          .
          <volume>3420415</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Giuste</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Naren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Isgut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gupte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Explainable artificial intelligence methods in combating pandemics: A systematic review</article-title>
          ,
          <source>IEEE Reviews in Biomedical Engineering</source>
          <volume>16</volume>
          (
          <year>2023</year>
          )
          <fpage>5</fpage>
          -
          <lpage>21</lpage>
          . doi:
          <volume>10</volume>
          .1109/RBME.
          <year>2022</year>
          .
          <volume>3185953</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ahnaf Alavee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Hasnayen</given-names>
            <surname>Zillanee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mostakim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uddin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Silva</given-names>
            <surname>Alvarado</surname>
          </string-name>
          , I. de la Torre Diez, I. Ashraf,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Abdus Samad, Enhancing early detection of diabetic retinopathy through the integration of deep learning models and explainable artificial intelligence</article-title>
          ,
          <source>IEEE Access 12</source>
          (
          <year>2024</year>
          )
          <fpage>73950</fpage>
          -
          <lpage>73969</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2024</year>
          .
          <volume>3405570</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sadeghi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Alizadehsani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>CIFCI</surname>
          </string-name>
          , S. Kausar,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rehman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mahanta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Bora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almasri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Alkhawaldeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Alatas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shoeibi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Moosaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hladík</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nahavandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Pardalos</surname>
          </string-name>
          ,
          <article-title>A review of explainable artificial intelligence in healthcare</article-title>
          ,
          <source>Computers and Electrical Engineering</source>
          <volume>118</volume>
          (
          <year>2024</year>
          )
          <article-title>109370</article-title>
          . URL: https: //www.sciencedirect.com/science/article/pii/S0045790624002982. doi:https://doi.org/ 10.1016/j.compeleceng.
          <year>2024</year>
          .
          <volume>109370</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kirmani</surname>
          </string-name>
          , G. Kour,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mohd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sheikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Maqbool</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Wani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Wani</surname>
          </string-name>
          ,
          <article-title>Biomedical semantic text summarizer</article-title>
          ,
          <source>BMC Bioinformatics 25</source>
          (
          <year>2024</year>
          )
          <article-title>152</article-title>
          . doi:
          <volume>10</volume>
          .1186/ s12859-024-05712-x.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>I. A. E.</given-names>
            <surname>Madi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Redjdal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bouaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Séroussi</surname>
          </string-name>
          ,
          <article-title>Exploring explainable ai techniques for text classification in healthcare: A scoping review</article-title>
          ,
          <source>Studies in Health Technology and Informatics</source>
          <volume>316</volume>
          (
          <year>2024</year>
          )
          <fpage>846</fpage>
          -
          <lpage>850</lpage>
          . doi:
          <volume>10</volume>
          .3233/shti240544.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>MacAvaney</surname>
          </string-name>
          , et al.,
          <article-title>Ontology-aware clinical abstractive summarization</article-title>
          ,
          <source>in: SIGIR'19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1013</fpage>
          -
          <lpage>1016</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>