<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Italian Conference on Big Data and Data Science, September</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>marization via Medical Entity Recognition and Generative AI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giuseppe Riccio</string-name>
          <email>giuseppe.riccio9@studenti.unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Romano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andriy Korsun</string-name>
          <email>a.korsun@studenti.unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Cirillo</string-name>
          <email>michele.cirillo2@studenti.unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Postiglione</string-name>
          <email>marco.postiglione@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio La Gatta</string-name>
          <email>valerio.lagatta@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonino Ferraro</string-name>
          <email>antonino.ferraro@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Galli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincenzo Moscato</string-name>
          <email>vmoscato@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Named Entity Recognition, Entity Linking, Relation Extraction, Summarization, Generative AI</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BIG DATA CINI National LAB - Node University of Naples ”Federico II”</institution>
          ,
          <addr-line>Naples, 80125</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Naples ”Federico II”</institution>
          ,
          <addr-line>Via Claudio 21, Naples, 80125</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>1</fpage>
      <lpage>13</lpage>
      <abstract>
        <p>This paper presents a fully automated approach for extracting value from content that lies hidden in Electronic Health Records (EHRs) using Large Language Models (LLMs) and Natural Language Processing (NLP) techniques, such as Named Entity Recognition (NER) and Entity Linking (L). In particular, the state-of-the-art approaches used to solve this task sufer from problems related to poor automation, given the laborious process of fine-tuning the models used and the dificult interpretation of the results obtained from them. The solution proposed in this work, on the other hand, aims to show the potential of NLP and generative AI to extract the relevant medical concepts contained within EHRs and generate a summary of the entire clinical history of each patient to construct a simple and intuitive dashboard that supports medical personnel with relevant medical information and useful analytics in order to diagnose and make decisions regarding the clinical condition of a patient.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        The exponential increase of complex aggregated data in the healthcare sector and beyond has
made understanding such data by medical professionals a challenging task. Extracting relevant
information from Electronic Health Records [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], such as clinical notes, represents a significant
challenge that requires advanced solutions based on the potential of Big Data and Natural
Language Processing (NLP). In this article, we review the current approaches proposed in
the scientific literature and present a possible solution that takes advantage of Named Entity
LGOBE
CEUR
Workshop
Proceedings
Recognition (NER), Entity Linking (L), Relation Extraction (RE) and text synthesis techniques
such as Large Languages Models (LLMs).
      </p>
      <p>
        In the scientific literature, several approaches have been proposed for the automatic extraction
and synthesis of information from clinical records. One of the main approaches used is Named
Entity Recognition (NER) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], that focuses on the identification and extraction of relevant
entities, which in this case could include diseases, drugs, medical procedures and symptoms
within clinical texts. The use of clustering techniques [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is common to organise and categorise
extracted data, e.g. by applying clustering algorithms to group together notes dealing with
similar topics. Furthermore, text synthesis techniques are used through the use of Generative
Language Models [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which enable the generation of coherent and contextually appropriate
summaries based on the extracted data, for example, reports on a patient’s medical history,
providing a clear and concise picture of medical conditions, prescribed drugs and relevant
procedures performed.
      </p>
      <p>This paper presents a summary of the final project conducted as part of the Big Data
Engineering course ofered in the Master’s Degree program in Computer Science at the University
of Naples Federico II. The project aimed to address the challenges of the Big Data domain, with
a specific focus on data management, processing, analysis, report generation, and protection.
Our proposal is based on the combination of advanced techniques previously discussed, such as
NER and LLMs.</p>
      <p>To evaluate the efectiveness of our solution, we employed the MIMIC III dataset as our data
source. This dataset is widely recognized for its extensive coverage, representativeness, and
richness of clinical information. Through the application of our approach to this dataset, we
demonstrate the information extraction and synthesis process, presenting the obtained results
and their validity in the clinical context.</p>
      <p>With this article, we fill a significant gap in the scientific literature by making a relevant
contribution to the development of advanced approaches for the automatic extraction and
synthesis of concepts from medical records, exploiting the potential of Big Data and LLMs to
provide a comprehensive view of a patient’s medical condition over time.</p>
      <p>The results obtained from this study are valuable both for the academic community and the
industry, as they ofer a solid foundation for further research and advancements in the field of
clinical data management and data engineering.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related Work</title>
      <p>
        In other fields than biomedical, several annotation interfaces have been developed for popular
Natural Language Processing (NLP) tasks, such as Named Entity Recognition (NER), Entity
Linking (EL), Relation Extraction (RE), Entity Normalization, Dependency Parsing, Chunking
and so on. Among the available options, open-source tools such as BRAT (Stenetorp et al.,
2012 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) have gained popularity. BRAT not only facilitates the management, monitoring and
collection of annotated document corpuses, but also supports general annotation tasks. Another
tool, Prodigy1, is a commercial product that ofers a modern annotation method for creating
training and evaluation data for machine learning models. Although this tool can use various
      </p>
      <sec id="sec-3-1">
        <title>1Documentation available at this site: https://prodi.gy/docs</title>
        <p>
          models to suggest entities, being based on the well-known NLP SpaCy library, it lacks automated
integration with existing biomedical NER+L systems. Regarding biomedical NER+L, previous
scientific research has introduced tools such as MetaMAP (Aronson, 2001 [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]) and CTakes
(Savova et al., 2010 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]). These tools allow users to inspect recognized entities but do not provide
mechanisms to correct and refine concepts or specify additional annotations based on specific
research areas. Another tool called SemEHR (Wu et al., 2018 [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]) focuses on biomedical NER+L
but difers in its approach from previous tools. Indeed, SemEHR allows for the incorporation of
customized preprocessing and postprocessing steps and supports research-specific use cases.
However, it does not directly improve the NER+L model through an interface, but treats the
provided NER+L model as a black-box model, with no possibility of changing the recognized
entities or obtaining more in-depth meta-information on those returned.
        </p>
        <p>
          Regarding the summarization task, there are two main approaches in the literature: the
ifrst, called extractive summarization, involves the generated summary being composed of
sentences extracted from the text provided as input based on a metric of importance of those
sentences in the context of the text. The second approach, called abstractive summarization,
involves extracting words within the text and reprocessing them to compose semantically
related sentences. Regarding extractive summarization, several solutions have been proposed
involving the use of neural networks, in which the problem is formulated as a classification
task and networks composed of encoders and decoders are used (Cheng and Lapata, 2016 [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ];
Nallapati et al., 2016 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]) or pre-trained language models (Egonmwan and Chali, 2019 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]; Liu
and Lapata, 2019 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]). In contrast, the state of the art in abstractive summarization involves
numerous approaches, the most popular of which is the use of pre-trained encoder-decoder
transformer models on a masked pre-training input target, the most popular of which are MASS
(Song et al., 2019 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]), UniLM (Dong et al., 2019 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]), and BART (Lewis et al., 2019 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]).
        </p>
        <p>However, the approaches just mentioned require manual work to classify each document
and the concepts it contains, and it is impractical to manually annotate large datasets such as
those of patient notes. In this paper we try to explore, therefore, not only the development of a
simple interface for annotating clinical texts, through NER+L models but also the generation
of complete summaries of a patient’s clinical history using LLMs, which allows to provide
these summaries starting from the clinical notes treated in an appropriate way in a rapid and
completely automated way with no need to fine tune a model as required by other approaches.
Furthermore, through the integration of the entities extracted from the NER+L+RE and with
the patient’s summary, it is possible to provide, through a simple and intuitive dashboard, a
series of analytics that support the diagnoses and decisions to be made with respect to an ill
patient by competent medical personnel.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Methodological Workflow</title>
      <p>Within our study, therefore, our goal was to develop a solution for the automatic summarization
and report generation of clinical notes using, as previously written, a combination of Named
Entity Recognition (NER), Linking (L), Relation Extraction (RE) and Language Models (LLMs).
In order to achieve this, we followed a detailed workflow as follows: (Figure 1)</p>
      <sec id="sec-4-1">
        <title>3.1. Data collection</title>
        <p>
          We used the anonymised clinical database provided by MIMIC-III [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. In particular,
we extracted information from the tables ”PATIENTS” and ”NOTEEVENTS”. This data collection
phase forms the basis of our study, although further processing is necessary to improve the
data quality.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Data extraction</title>
        <p>We conducted a data analysis to understand the distribution and composition of the clinical
notes. This analysis helped us decide which fields to extract. During the data extraction process,
we applied additional filters according to the limitations imposed by the selected data storage.
In particular, we set a maximum limit of 10,000 tokens for the sum of the documents. We
performed additional filtering to select at least 2 documents per patient and removed duplicate
rows. Moreover, we extracted the demographic information of the patients from the ”PATIENTS”
table and merged it with the clinical notes dataframe to obtain a final dataframe upon which to
perform subsequent project operations. The final data was stored in the selected data storage.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Data preprocessing</title>
        <p>We selected a sample of the previously extracted data in order to make the further processes
more eficient. Next, we applied the Data Preprocessing step to the sampled clinical notes. This
operation included the extraction of relevant sections from the clinical notes (Section Extraction)
and the application of textual preprocessing techniques, such as stopwords elimination and
lemmatization, to the text of the clinical notes.</p>
      </sec>
      <sec id="sec-4-4">
        <title>3.4. Summarization</title>
        <p>To generate the summaries of the clinical notes, we used a model based on Large Language
Models (LLMs). The input for the model was structured according to the following prompt:
• System: e.g for summary: ”You are a formal medical assistant specialising in the summary
of a patient’s clinical notes” while for clinical trend: ”You are a clinical trend extractor of
a patient ’s clinical notes.”
• User: an experimental approach:
 =  +  + 
(1)
Where:
– P: Prompt;
– E: Explanation: explanation of what is wanted from the model;
– T: Text: text that the model must deal with and summarise, in this case the clinical
notes;
– S: Show the type of the task: demonstration of an example of output useful for the
model;
Explanation
Generate a complete but concise (max 100 words) and informative summary that
focuses only on the unique patient, that is always the same throughout the notes in
input, and his medical history, current condition, and relevant details, starting from
his current conditions backwards; The summary must be accurate and avoid
unnecessary repetition, avoid details of prescriptions, doses and other subjects
besides the specific patient; If there are dates, enter the year, month and the day; Do
not include doctors' names; It must be easy for a doctor to understand, the tone is
formal; The summary always starts with "The patient is a:" ".</p>
        <p>Text
here goes the sequence of clinical notes representing the medical history</p>
        <p>Show
The patient currently is a age-year-old gender who presented with chief complaint
and has a medical history of relevant medical conditions. The patient was involved
in incident/accident description, which led to specific injuries/traumas; The
postoperative course involved relevant procedures performed on anatomical
locations involved; The patient is currently prescribed medications for specific
purposes;</p>
        <p>Explanation
Identify in 1 word the clinical trend like this: extract a single word that accurately
represents the clinical trend observed in the patient's notes, considering every detail
and keyword; For the word, use the general indicators: 'Improvement', 'Stable',
'Worsening' to describe the trend; Otherwise identify in 1 word: "Dead" if the
patient from his notes is dead.</p>
        <p>Text
here goes the sequence of clinical notes representing the medical history</p>
        <p>Show</p>
        <p>Improvement
(a) Summary
(b) Clinical trend</p>
        <p>We applied the same structure for both the general summary of the clinical notes (Summary)
and the identification of clinical trends (Clinical Trend).</p>
        <p>Summary: the aim of the summary is to provide an overview of the patient’s condition,
treatments carried out, main diagnoses or other key points in the clinical notes. Figure 2(a)
shows an example of summarization prompts.</p>
        <p>Clinical trend: refers to patterns or changes observed in a patient’s clinical picture over
time. Figure 2(b) shows an example of prompt used to infer the medical history trend.</p>
      </sec>
      <sec id="sec-4-5">
        <title>3.5. Named Entity Recognition + Linking + Relation Extraction</title>
        <p>From the preprocessed data, we applied the Named Entity Recognition (NER) and Entity Linking
(L) process to extract medical entities from the clinical notes and link them to external knowledge
databases. In addition, we performed Relation Extraction (RE) to identify possible relationships
between entities. This information was used to create a Knowledge Base for graph analysis, in
which diagnostic analyses were performed.</p>
      </sec>
      <sec id="sec-4-6">
        <title>3.6. Analytics and Report</title>
        <p>We summarized the results obtained from the clinical note summary process and the Named
Entity Recognition + Entity Linking + Relation Extraction process. These results were presented
in a final report documenting the main results obtained within our study.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <sec id="sec-5-1">
        <title>4.1. Implementation details</title>
        <p>
          The techniques seen above are applied by randomly selecting 20 patients from the
MIMICIII dataset [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]; in particular, only the ”Discharge summary” of patients selected from the
NOTEEVENTS table is considered2.
        </p>
        <p>
          Then, through the MedSpaCy [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] module, the extraction of the main sections, i.e., those
most present, from all patients’ clinical notes is performed. MedSpaCy was chosen because that
library is already pre-trained to recognize sections present in clinical texts. Thus, for the case
under consideration, the following sections are selected: ”chief complaint,” ”history of present
illness,” ”past medical history,” ”discharge medications,” ”brief hospital course” and ”discharge
diagnoses”. In the clinical notes with the extracted sections, a following stage of stopwords
elimination and lemmatization is carried out through the NLTK3 module using stopwords of
the English language, the clinical notes being in that language, and WordNet as a lemmatizer.
        </p>
        <p>To perform the summarization task we chose to use as LLM the GPT-3.5-turbo 4 model
provided by OpenAI, this model, in fact, is the one that starting from a prompt and the clinical
notes of the incoming patient manages to return a short, but at the same time complete, summary
containing all the main information regarding the patient’s medical history. In addition to the
patient summary, the patient’s clinical trend is also generated using the same model, which,
based on the evolution of his clinical history shown in the notes, provides an indication of the
patient’s current status.</p>
        <p>
          For the NER task it was decided to use the MedCAT [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] model, this model was found
to be the one with the best performance for the requested task being pre-trained for natural
language processing in the medical field. In particular, it is very useful for extracting information
from Electronic Health Records (EHR) (NER phase) and linking them to biomedical ontologies
such as SNOMED-CT and UMLS (Entity Linking phase). Since MedCAT is only a model that
correctly extracts and labels entities, it is necessary to provide MedCAT with a knowledge
base from which to draw this information. For the case in question, it was decided to use the
MedMentions [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] library, which contains a corpus of biomedical documents annotated with
mentions of entities belonging to UMLS. This corpus contains about 35,000 medical concepts
and its MetaCAT model for meta annotations was built on a sample from MIMIC-III. An example
of a clinical note annotated with MedCAT using MedMentions is shown in Figure 3.
        </p>
        <p>Finally, the concepts extracted from the NER+L phase must be interconnected through the
relationships extracted with the Relation Extraction phase, also in this case these relations
are obtained using the GPT-3.5-turbo model. Given their extremely complex nature, these
relationships are particularly suitable to be stored via a graph database such as Neo4J5.
2The complete code used to carry out the experiments reported in the following article is available in the following
GitHub repository: https://github.com/giuseppericcio/HealthcareSummarizationMIMIC
3Documentation available at the site: https://www.nltk.org/
4More details on models provided by OpenAI: https://platform.openai.com/docs/models/gpt-3-5
5Documentation available at the site: https://neo4j.com/docs/</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Dashboard and Analytics</title>
        <p>Thanks to the entities extracted from the NER+L+RE phase and the summaries generated by the
LLM, it is possible to build a dashboard that supports the medical personnel during the diagnosis
and the choice of therapies to be carried out on a patient in order to cure his diseases. In this
instance, it was decided to use the Streamlit6 library for the construction of the dashboard
directly in Python. This choice is dictated by the simplicity of creating a dashboard using
this library which allows various efective views of the analytics that can be performed on
biomedical concepts recognized by NER+L+RE. In Figure 4 is possible to see some examples of
summaries generated from some clinical notes of MIMIC-III patients, through these summaries
a physician can understand all the diseases and procedures with respect to that patient in a fast
way than reading all of his clinical notes. From Figure 4(b), a hypothetical flaw in our workflow
emerges concerning the use of male gender instead of female in the summary. This discrepancy
is probably due to errors resulting from possible hallucinations of the LLMs model. We propose
to resolve this issue by extending the previously used prompt or by adopting a more advanced
model such as GPT 4.</p>
        <sec id="sec-5-2-1">
          <title>6Documentation available at the site: https://docs.streamlit.io/</title>
          <p>4.2.1. Analytic 1: Lists of concepts extracted
First of all, through the dashboard, it is possible to view lists of drugs, symptoms and diagnostic
procedures related to a specific patient, as shown in Figure 5. Through this visualization, medical
personnel can immediately understand what the patient has been subjected to without having
to read all his medical records.
4.2.2. Analytic 2: Diseases of a patient
Using the relations extracted from the NER+L+RE phase is possible to visualize some
interesting analytics. Through the concepts stored in Neo4J, which has also been integrated into
Streamlit via the py2neo and streamlit-agraph libraries. As shown in Figure 6(a), it is possible to
efectively display all the diseases associated with a patient. For example, the patient taken into
consideration presented ”Hypoxia” in all the medical records associated with him, therefore, it
could be deduced that he sufered from it chronically.
4.2.3. Analytic 3: Medical concepts related to a disease of a patient
With reference to the previous analytic, we can now explore all the diagnostic procedures,
drugs and other treatments performed on the patient to cure the ”Hypoxia” disease, as shown
in Figure 6(b), in order to facilitate physicians and nurses understand what has already been
done and what needs to be done later on to that patient.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion</title>
      <p>Our solution ofers numerous benefits, such as automating the clinical note synthesis process,
improving productivity and reducing analysis time for healthcare professionals. Through the
use of advanced techniques such as NER+L+RE, we are able to identify medical entities and
the relationships between them, providing a solid basis for analysing and interpreting clinical
information.</p>
      <p>However, the automatic extraction and synthesis of information from clinical records presents
significant and sensitive challenges. The management of sensitive data, in compliance with
regulations such as the General Data Protection Regulation (GDPR), the Health Insurance
(a) Diseases of the patient
(b) Procedures, drugs and so on related to a
specific disease
Portability and Accountability Act (HIPAA), and the Artificial Intelligence Act, requires special
attention to privacy and the protection of sensitive personal data. Therefore, it is essential to
ensure security and compliance with guidelines regarding data access, storage and disposal.</p>
      <p>To address these ethical and regulatory issues, robust security measures must be implemented
to protect patients’ personal data and ensure transparency in the use of clinical information.
Furthermore, it is crucial to inform patients about the processing of their data and the generative
nature of the results obtained, emphasising that the system does not replace the work of doctors.
Adherence to ethical standards is of paramount importance to maintain patients’ trust and to
ensure responsible and safe use of clinical information. For example, the automatic generation
of summaries raises ethical issues regarding the accurate interpretation of clinical information,
so it is necessary to ensure the reliability of the generated results and assess the quality of the
summaries through further research and validation.</p>
      <p>Another disadvantage comes from the tools and limitations of the models used. The specialised
medical language used in clinical registries, with abbreviations and technical terminology,
requires the use of additional resources, such as medical dictionaries or abbreviation recognition
systems, in order to overcome the challenges of information interpretation and extraction.
Furthermore, the size of clinical note datasets requires the use of eficient text processing models
in order to handle large volumes of data.</p>
      <p>Finally, our work represents a step towards automation and optimization of clinical note
analysis, but further studies and collaborations are needed to improve the accuracy, reliability
and adherence to ethical standards of our solutions.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusion and Future Work</title>
      <p>Our work has developed a solution for the automated summarization of clinical notes using
NER+L+RE and LLM techniques, providing fast decision support for healthcare professionals
towards patients. The preliminary results obtained in this paper were submitted to a committee
of domain experts for review, who upon initial analysis evaluated the work positively. However,
there are many possibilities for further development and future work arising from this project.
Some ideas include:
• Predictive analytics: Expanding our solution to include predictive analytics models that
can provide estimates of patients’ future conditions, such as the likelihood of developing
certain diseases or response to certain treatments;
• Patient profiling : Create a comprehensive overview of all patients treated, allowing
clinicians to identify high-risk patients. This would require unsupervised data analysis,
such as using clustering algorithms to group patients according to common characteristics;
• Interactivity: Increased interactivity through human body diagrams that display diseased
or clinically afected body parts with a brief summary of the problem;
• Integration: Expand our solution to be easily integrated with databases from diferent
hospitals, allowing healthcare professionals to use the system with their own data;
• Interpretability: Improve the transparency and interpretability of the system by
providing clear explanations of the forecasts and recommendations generated. This would
help physicians understand the reasons behind the results and have confidence in the
information provided.
• Q/A (Question/Answer): Implement a Q/A interface that allows doctors and patients to
interact directly with the system, asking specific questions and getting precise answers
based on the data in the database.</p>
      <p>In conclusion, the future goal is to continue to develop solutions that improve the eficiency
(currently the proposed solution on 20 patients takes an average time of 298 seconds) and
accuracy of clinical note analysis through systematic and more formal approaches.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Charles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Furukawa</surname>
          </string-name>
          ,
          <article-title>Adoption of electronic health record systems among u.s. non-federal acute care hospitals</article-title>
          ,
          <source>ONC Data Brief No. 9</source>
          (
          <issue>2013</issue>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Soomro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Banbhrani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shaikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Raj</surname>
          </string-name>
          ,
          <article-title>Bio-ner: Biomedical named entity recognition using rule-based and statistical learners</article-title>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Loftus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shickel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Balch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tighe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Abbott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fazzone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rozowsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Ozrazgat</given-names>
            <surname>Baslanti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Berceli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Efron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Moorman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rashidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Upchurch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bihorac</surname>
          </string-name>
          ,
          <article-title>Phenotype clustering in health care: A narrative review for clinicians</article-title>
          ,
          <source>Frontiers in Artificial Intelligence</source>
          <volume>5</volume>
          (
          <year>2022</year>
          )
          <article-title>842306</article-title>
          . doi:
          <volume>10</volume>
          .3389/frai.
          <year>2022</year>
          .
          <volume>842306</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M. K.</given-names>
            <surname>Rohil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Magotra</surname>
          </string-name>
          ,
          <article-title>An exploratory study of automatic text summarization in biomedical and healthcare domain</article-title>
          ,
          <source>Healthcare Analytics</source>
          <volume>2</volume>
          (
          <year>2022</year>
          )
          <article-title>100058</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.health.
          <year>2022</year>
          .
          <volume>100058</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Stenetorp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          , G. Topic,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ohta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ananiadou</surname>
          </string-name>
          , J. Tsujii,
          <article-title>brat: a web-based tool for nlp-assisted text annotation, The 3th Conference of the European Chapter of the Association for Computational Linguistics</article-title>
          ; Avignon, France (
          <year>2012</year>
          )
          <fpage>102</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Aronson</surname>
          </string-name>
          ,
          <article-title>Efective mapping of biomedical text to the umls metathesaurus: The metamap program</article-title>
          ,
          <source>Proceedings / AMIA ... Annual Symposium. AMIA Symposium</source>
          <year>2001</year>
          (
          <year>2001</year>
          )
          <fpage>17</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Masanz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ogren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kipper-Schuler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chute</surname>
          </string-name>
          ,
          <article-title>Mayo clinical text analysis and knowledge extraction system (ctakes): Architecture, component evaluation and applications</article-title>
          ,
          <source>Journal of the American Medical Informatics Association : JAMIA</source>
          <volume>17</volume>
          (
          <year>2010</year>
          )
          <fpage>507</fpage>
          -
          <lpage>13</lpage>
          . doi:
          <volume>10</volume>
          .1136/jamia.
          <year>2009</year>
          .
          <volume>001560</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wu</surname>
          </string-name>
          , G. Toti,
          <string-name>
            <given-names>K.</given-names>
            <surname>Morley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ibrahim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Folarin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Jackson</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kartoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gorrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Broadbent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stewart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dobson</surname>
          </string-name>
          ,
          <article-title>Semehr: A general-purpose semantic search system to surface semantic data from clinical notes for tailored care, trial recruitment, and clinical research</article-title>
          ,
          <source>Journal of the American Medical Informatics Association</source>
          <volume>25</volume>
          (
          <year>2018</year>
          )
          <article-title>160</article-title>
          . doi:
          <volume>10</volume>
          .1093/jamia/ocx160.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cheng</surname>
          </string-name>
          , M. Lapata,
          <article-title>Neural summarization by extracting sentences and words</article-title>
          ,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P16</fpage>
          - 1046.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Nallapati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>Summarunner:</surname>
          </string-name>
          <article-title>A recurrent neural network based sequence model for extractive summarization of documents</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>31</volume>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1609/aaai.v31i1.
          <fpage>10958</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Egonmwan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chali</surname>
          </string-name>
          ,
          <article-title>Transformer-based model for single documents neural summarization</article-title>
          ,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          - 5607.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lapata</surname>
          </string-name>
          ,
          <article-title>Text summarization with pretrained encoders</article-title>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lu</surname>
          </string-name>
          , T.-Y. Liu,
          <article-title>Mass: Masked sequence to sequence pre-training for language generation</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1905</year>
          .02450.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , H.-W. Hon,
          <article-title>Unified language model pre-training for natural language understanding and generation</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1905</year>
          .03197.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghazvininejad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          , L. Zettlemoyer, Bart:
          <article-title>Denoising sequence-to-sequence pre-training for natural language generation, translation</article-title>
          , and comprehension,
          <year>2019</year>
          . arXiv:
          <year>1910</year>
          .13461.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , T. Pollard, M. Roger, ”
          <article-title>mimic-iii clinical database”</article-title>
          <source>(version 1.4)</source>
          ,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          . 13026/C2XW26.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , T. Pollard,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shen</surname>
          </string-name>
          , L.-w. Lehman,
          <string-name>
            <given-names>M.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghassemi</surname>
          </string-name>
          , B. Moody, P. Szolovits,
          <string-name>
            <given-names>L.</given-names>
            <surname>Celi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mark</surname>
          </string-name>
          ,
          <article-title>Mimic-iii, a freely accessible critical care database</article-title>
          ,
          <source>Scientific Data</source>
          <volume>3</volume>
          (
          <year>2016</year>
          )
          <article-title>160035</article-title>
          . doi:
          <volume>10</volume>
          .1038/sdata.
          <year>2016</year>
          .
          <volume>35</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Goldberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Amaral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Glass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Havlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hausdorg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mietus</surname>
          </string-name>
          , G. Moody, C.
          <article-title>-</article-title>
          K. Peng,
          <string-name>
            <given-names>H.</given-names>
            <surname>Stanley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Physiobank</surname>
          </string-name>
          ,
          <article-title>Components of a new research resource for complex physiologic signals</article-title>
          ,
          <source>PhysioNet</source>
          <volume>101</volume>
          (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>H.</given-names>
            <surname>Eyre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Peterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Alba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Box</surname>
          </string-name>
          , S. DuVall,
          <string-name>
            <given-names>O.</given-names>
            <surname>Patterson</surname>
          </string-name>
          ,
          <article-title>Launching into clinical space with medspacy: a new clinical text processing toolkit in python</article-title>
          ,
          <source>AMIA ... Annual Symposium proceedings. AMIA Symposium</source>
          <year>2021</year>
          (
          <year>2022</year>
          )
          <fpage>438</fpage>
          -
          <lpage>447</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kraljevic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mascio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Roguski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Folarin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bendayan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dobson</surname>
          </string-name>
          , Medcat - medical
          <source>concept annotation tool</source>
          ,
          <year>2019</year>
          . arXiv:
          <year>1912</year>
          .10166.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Medmentions: A large biomedical corpus annotated with umls concepts</article-title>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>