<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Diagnosis and Characterisation of Prostate Cancer in Pathology Reports in Spanish</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rosa M. Montañés-Salas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergio Gracia-Borobia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>María de la Vega Rodrigálvarez-Chamarro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ángel Borque-Fernando</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patricia A. Guerrero-Ochoa</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandro Camón-Fernández</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Alfaro-Torres</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Isabel Marquina-Ibáñez</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sofia Hakim-Alonso</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis M. Esteban</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael del-Hoyo-Alonso</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aragon Institute of Technology (ITA), María de Luna</institution>
          ,
          <addr-line>7-8, 50018 Zaragoza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Applied Mathematics, Escuela Universitaria Politécnica de La Almunia, Universidad de Zaragoza</institution>
          ,
          <addr-line>50100 Zaragoza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Pathology, Miguel Servet University Hospital (GIIS071-uro-servet)</institution>
          ,
          <addr-line>50009 Zaragoza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Department of Urology, Miguel Servet University Hospital (GIIS071-uro-servet)</institution>
          ,
          <addr-line>50009 Zaragoza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Health Research Institute of Aragon Foundation (GIIS071-uro-servet)</institution>
          ,
          <addr-line>50009 Zaragoza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Prostate cancer is a prevalent disease worldwide, with early diagnosis enabling better prognosis. Natural language processing (NLP) techniques show promise in extracting information from electronic health records to support clinical decision-making. This paper presents an NLP approach to detect and characterise prostate cancer (PCa) diagnoses from Spanish pathology reports. A combination of lexical-morphological analysis, rule-based techniques and transformer models is used to identify PCa, Gleason scores, procedures, organs and other markers. The system achieves near 96% agreement in detecting cancer diagnoses compared to expert annotation.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Medical Natural Language Processing</kwd>
        <kwd>Prostate Cancer</kwd>
        <kwd>Pathology Reports</kwd>
        <kwd>Information Extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Prostate cancer (PCa) was the fourth most diagnosed
cancer worldwide in 2022. In Spain, in 2023, there have been
reported between 33,000-34,000 new cases diagnosed, with
a 5-year prevalence of more than 140,000 cases, making it
the leading cancer in terms of incidence among the male
population, as reported by the Spanish Cancer Association
in 2023 1. Early detection of PCa enables treatment at initial
stage, resulting in higher cure rates and reduced side efects
of aggressive treatments, as well as lower healthcare costs.
In this context, the use of advanced Artificial Intelligence
(AI) techniques presents itself as a promising tool to
support clinical decision-making systems and predictive model
research [
        <xref ref-type="bibr" rid="ref1 ref25">1</xref>
        ]. Specifically, Natural Language Processing can
play a decisive role by facilitating information retrieval from
non-structured sources and opening the range of processing
techniques [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        The acquisition of large volumes of patient data for
research purposes has become feasible today, thanks to the
digitization of healthcare systems and systematic recording
of medical procedures. However, there are still limitations
stemming from the non-uniformity of information systems,
the diverse repositories for clinical analyses, radiological
or pathological reports, or other Electronic Health Records
(EHRs), and the assessment and follow-up of patients
conducted by diferent healthcare professionals. Therefore, the
sources and data are significantly heterogeneous,
including a substantial amount of textual information containing
valuable clinical knowledge provided by experts in the field,
which allows for precise and accurate diagnoses. Moreover,
privacy concerns have to be taken into account when
dealing with patient sensitive data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The research presented here is part of the
AI4HealthyAging project, whose mission is to
leverage distributed AI technologies for early diagnosis and
treatment of diseases that are highly prevalent in the ageing
population. Within this project, several work packages are
organized regarding various diseases such as Parkinson,
sarcopenia, deafness or cancer. The use case related to the
diagnostic management of prevalent cancers in the elderly
is focused on prostate and colon cancers. Particularly, the
overarching goal for prostate cancer is to develop decision
support and risk interpretation tools based on the biological
footprint present in patient EHRs, histological preparations,
and radiological images, thus enhancing diagnosis through
the use of hybrid data.</p>
      <p>In this article, we present our experience in analysing,
extracting and structuring information using Natural
Language Processing techniques applied to pathology reports of
prostate cancer in Spanish. The main challenge presented is
the unavailability of a truly reliable and medically consistent
labeled dataset in Spanish from which to incorporate new
clinical features for the development of advanced hybrid
predictive models. Therefore, the first stages of this project
consisted of the development of a working methodology to
retrieve relevant data and implement an NLP-based system
that would enable to eficiently detect and characterise
cancer diagnoses based on pathologists’ reports. The proposed
approach has facilitated the compilation of a comprehensive
and validated dataset, providing both, a final/clear
diagnosis and several features of interest, for the development of
explainable risk prediction and prognosis machine learning
models fueled by multiple information sources.</p>
      <p>The paper is organised as follows: a contextual
introduction that delineates the research setting has been outlined,
followed by a related work exploration. Section 3
encompasses the materials and methods integral to underpinning
the system’s development. Subsequently, the attained
results are discussed upon finishing with the conclusions of
the work and prospective areas for future development and
research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        Natural Language Processing applied to the biomedical
domain has witnessed significant growth and innovation in
recent years, driven by the increasing availability of
largescale healthcare data and advancements in NLP techniques.
Electronic health records have emerged as a significant
source of information for the detection and diagnosis of
several diseases, including cancer. With the incorporation
of rich textual data, natural language processing has been
extensively applied to these records. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] systematically
reviewed the applications of NLP in sifting through EHRs,
highlighting its potential in detecting chronic diseases and
signs of various cancer types, respectively, emphasizing
the challenges of data heterogeneity and the significance of
domain-specific annotations. This is also shown in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], where information was extracted from free-text
pathology reports related to breast and lung cancer and colorectal
cancer, respectively, both using expert annotated textual
data.
      </p>
      <p>
        In the realm of prostate cancer, a niche yet rapidly
growing area of research focuses on leveraging NLP techniques
for diagnostic purposes. DiBello et al. demonstrated that
NLP can accurately identify metastatic PCa by searching
unstructured text in medical records such as pathology,
radiology and clinic notes. Thomas et al. validated an NLP
program to accurately identify patients with prostate cancer
and extract relevant information from pathology reports.
Some approaches also underscore the critical role of
domainspecific knowledge in curating and understanding the
specialized terminology present in such reports [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
Additionally, the extraction of such domain-specific knowledge helps
improve unimodal models within PCa diagnosis: Morote
et al. assessed the ability of microscopic findings in prostate
biopsies to improve the prediction of clinically significant
prostate cancer using numerical-only data, and Khosravi
et al. developed an AI based model for PCa diagnosis using
magnetic resonance images labelled with manually-assigned
histopathology information.
      </p>
      <p>
        A great variety of text-related machine learning models
have seen application in this domain: Breischneider et al.
developed an unsupervised rule-based ontology system for
feature extraction in free-form text obtained from clinical
reports. Yoon et al. dived deeper and demonstrates how graph
neural networks can be trained with in-context textual data
for multitask cancer labeling. Also, transformer-based
models, renowned for their prowess in NLP tasks, have been
introduced into the biomedical field. ClinicalBioBERT [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
and OncoBERT [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] showcased the utility of BERT and its
variants in comprehending medical narratives, aiming at
identifying signs of cancers. Their findings illuminate the
potential of fine-tuning such models with domain-specific
data to boost diagnostic performance.
      </p>
      <p>Despite these advancements, challenges persist in
achieving optimal performance for the detection and diagnosis
of diseases, particularly due to the inherent noise and
variability in EHRs and other medical reports. Moreover, the
multilingualism challenge is still an open issue, although
multiple eforts are being conducted to develop annotation
standards and AI systems in Spanish (Seda et al.;
MirandaEscalada et al.; Solarte-Pabón et al.), the scope of research
in the oncology field is still limited.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed approach</title>
      <p>This section outlines the proposed approach for the
development of the clinical report analysis system. The system
aims to extract relevant information from medical
documents and classify them according to their diagnosis or the
requirements proposed. It poses a working methodology
customizable to diferent types of medical text reports, in
which, starting from basic resources in the form of a textual
corpus and elementary terminology, diferent natural
language processing strategies are applied semi-automatically
to characterise the data set and obtain implicit information
useful for feeding other learning systems.</p>
      <p>The materials and methods presented here have been
developed by three independent teams in order to ensure
data privacy2 and broad applicability of the system. First,
the team responsible for accessing the healthcare databases
retrieves two sets of data. Thereafter, the text set is analysed
by the natural language processing team and its results are
combined with the medical data modelling team’s set for
validation purposes.</p>
      <sec id="sec-3-1">
        <title>3.1. Materials</title>
        <p>The initial study population in this work consists of all males
afiliated with Zaragoza II healthcare sector (392,177
individuals), who have been identified with at least one
ProstateSpecific Antigen (PSA) determination in the last 5 years, i.e.,
in the interval between 2017 and 2022. With these
characteristics, a total of 92,171 patients were identified. From this
initial population, those with at least one biopsy performed,
and consequently, with at least one pathology report
available in the system, were selected, resulting in a total of
10,563 patients of interest. All available data in the hospital
systems accountable for performing these procedures were
retrospectively collected since 1999, ultimately retrieving a
total of 24981 textual pathology reports.</p>
        <sec id="sec-3-1-1">
          <title>3.1.1. Data preparation</title>
          <p>The team responsible for accessing healthcare databases and
retrieving reports work independently of the other teams to
ensure the confidentiality and protection of sensitive data.
The former team has conducted a triple pseudonymization
process upon the pathology reports, compiling a document
database for natural language processing with two content
sections: the macroscopic description and the diagnosis
section. The macroscopic description provides concrete
2Following the Regulation (EU) 2016/679 of the European Parliament
and of the Council of 27 April 2016 on the protection of individuals with
regard to the processing of personal data and on the free movement of
such data (GDPR).
details regarding the tissue removal procedure, while in the
diagnosis section, pathologists determine the findings in the
extracted tissue samples and provide detailed information
they deem relevant for patient monitoring. Both, “diagnosis”
and “macro” sections are considered for the detection and
characterisation of PCa (referred to as “corpus” in figure 1).</p>
          <p>
            Along with this base corpus, a lightweight thesaurus was
built from the exploration of PCa-related concepts in
standard ontologies and classification schemas such as SNOMED
[
            <xref ref-type="bibr" rid="ref20">20</xref>
            ] and ICD10-codes [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ] in their Spanish versions. This
resource has been constructed by the NLP team with the
guidance of the medical staf at the Zaragoza II healthcare
sector in charge of reporting prostate cancer, with the aim of
simplifying the external knowledge integration,
implementing an easily adaptable tool and adjusting the linguistic
analysis to real-world usage by reproducing the communicative
registry employed in those healthcare reports, according to
the American College of Pathologist-CAP guidelines
actualized every 6 months by the specialist pathologist staf of
the Departament of Pathology. This dictionary is designed
as a simple hierarchy in which each concept or
characteristic to be extracted is paired with a carefully refined list
of expressions drawn from the standards mentioned above,
comprising the expert domain knowledge base of the
system.
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.1.2. Validation set</title>
          <p>To assess the results of the analysis system, an
independent dataset from morphology study repositories has been
retrieved under the same assumptions as the textual data
corresponding to the pathology reports, corresponding to
10297 patients with 22628 cases. These studies contain
numerical and categorical data, particularly a summary label
assigned by clinical staf during patient exams, following the
SNOMED nomenclature. A preliminary analysis of these
labels revealed a high dimensionality and variability of the
assigned classes. Additionally, excessively specific categories
were used, making it dificult to study PCa diagnosis
adequately. This diversity can be attributed to the complexity of
the standard used on a day-to-day basis and the subjectivity
of human criteria for assignment.</p>
          <p>An overview of the aforementioned resources is depicted
in the following figure 1.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Methods</title>
        <p>In the design of the clinical record information extraction
and labelling system depicted in figure 2, we aimed to follow
an iterative yet simple methodology: based on the set of
language resources outlined in the previous section (see 3.1)
and the application of diferent biomedical natural language
processing techniques, a comprehensive structured dataset
is built and refined in conjunction with the in-domain
knowledge. Both, dataset information and expert knowledge are
easily customized according to the specific needs of
subsequent machine learning models or expert requirements. In
our use case, a structured dataset of prostate cancer
diagnosis features is built and then validated by an independent
team, performing a limited number of feedback iterations
to improve the overall performance and guarantee the
generalization and customizable capabilities of the system.
The preliminary analysis phase consists of building
automatically a glossary of categorized terms and expressions that
are morphologically and semantically very close to the base
in-domain knowledge provided by the medical staf.
Therefore, a lightweight customized thesaurus is constructed,
through similarity searches over the corpus along with a set
of approximate regular morphological data patterns usually
found on cancer reports. These terms and expressions
facilitate the further identification of information patterns and
data in the textual records, aiding in the determination of
diagnoses and their context.</p>
        <p>From this, the available pathology reports are processed
and both the prostate cancer diagnosis, in terms of positive
or negative presence, and a number of additional features
of interest are inferred, building a rich dataset that can feed
other AI multimodal systems. To accomplish the detection
of diagnoses and the extraction of significant variables
employing the compiled lexicon and patterns, various strategies
are integrated into a three-step approach:</p>
        <p>Firstly, a content-based filtering is applied, discarding
empty or non-informative reports (i.e. many documents
consist of only “VER B” or similar texts). The second step
is based on a mixture of lexical-morphological analysis
supported by fuzzy matching and the integration of a
pretrained language model based on transformers. In this case,
we utilize biomedical language models in the Spanish
language, specifically the RoBERTa-base biomedical model that
has been already finetuned for the Named Entity
Recognition (NER) task on the Cantemist dataset for tumour
morphology extraction by Carrino et al.. The joint objective is
to extract cancer-related terms and expressions from the
text reports, which represent the relevant features pursued.
The third step consists of applying rule-based techniques
designed for the extraction of objective features. On the one
hand, a simple set of grammatical rules has been applied, i.e.
occurrence of certain constituents in the sentences, negation
detection and comparison of lexical-morphological analysis
with NER output. On the other hand, domain-specific rules
have been designed: a priority system between labels has
been established for the PCa diagnosis distinguishing from
positive cases to diferent gravity levels and not PCa as less
critical; specific rules to compute the final Gleason degrees
and groups when diferent values are retrieved; and warning
rules to analyse whether values extracted are incoherent.</p>
        <p>
          The detailed characteristics retrieved through the
described NLP techniques on the pathology reports of prostate
cancer are the following:
• Prostate Cancer diagnosis: a binary output (PCa+,
PCa-, for positive or negative Prostate Cancer
diagnosis, respectively) that includes an additional label
for uncertain cases in which the analysis and rule
processing throw opposing or empty results
(identiifed as Untagged). The latter serves to indicate the
need for a thorough review by an expert.
• Gleason Score (Sum and Group): the Gleason score,
often referred to as the Gleason grading system [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]
is a numerical scale used in the field of pathology
to assess the severity and aggressiveness of prostate
cancer glands under a microscope. It can be
expressed in several ways, generally consisting of two
numbers: a primary and a secondary cancer pattern,
i.e. the most common pattern of cancer glands seen,
and the next most common pattern, respectively
(tertiary is also extracted, but it is rarely mentioned).
The sum of the primary and secondary patterns
along with their corresponding group are identified
and conveniently computed. Whenever the Gleason
score is mentioned multiple times in a document
(i.e. in a pathology report describing a biopsy with
a list of inspected cilinders) all the possible
components are extracted but the gleason score associated
to that document is the most severe among the
multiple results. The pathology reports analysed in this
research show certain variability due to the
evolution of the Gleason grading system in the time range
considered, which have been redacted following the
recommendations and updates of the International
Society of Urological Pathology (ISUP).
• Type of Medical Procedure: each of the reports
analysed corresponds to a specific procedure conducted
on the patient. The considered procedures are:
Biopsy, Prostatectomy, Cystoprostatectomy,
Adenomectomy and TURP (Transurethral Resection of
the Prostate). An additional Untagged label is
considered in case none of the above could be detected.
        </p>
        <p>
          This result is treated as a multilabel output.
• Organ mentions: given that the diferent medical
procedures can afect several areas of the organism,
the mentioned organs on each document are also
extracted. The considered organs are: Prostate,
Seminal Vesicles, Lymph Nodes and Bladder. This result
is treated as a mutilabel output.
• Other informative markers: mentions to ASAP
(Atypical Small Acinar Proliferation); mentions to
PIN (Prostatic Intraepithelial Neoplasia); mentions
of inflammatory processes and atrophies; TNM
stage: standard classification for cancer staging, it
refers to Tumour, Nodes and Metastasis [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]; “DUC”
label related to mentions of ductal carcinoma; and
neoplastic morphology mentions extracted with the
NER model.
        </p>
        <sec id="sec-3-2-1">
          <title>3.2.2. Validation</title>
          <p>The modelling team, in collaboration with health
professionals, found that the SNOMED labels assigned in the
validation set (see section 3.1.2) are not very accurate for the
actual diagnosis of PCa. They have independently made
an adjustment to this set of labels, reducing it to a
threecategory classification (existence or non-existence of PCa
plus an ‘untagged’ label) through semi-automated mapping
and subsequent human validation by two experts.</p>
          <p>The intersection of the text corpus and the morphology
data over patient studies consists of 10249 patients and a
total of 17696 studies on which it is possible to compare
and validate PCa tagging results at document-level. The
inter-annotator agreement (IAA) is computed between the
NLP tags and the postprocessed human-assigned SNOMED
labels.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>The results reported in this section correspond to the final
results obtained after three cycles of compilation, execution
and validation of information extraction and classification
over the available pathology reports. The base outcome of
the system presented is the structured dataset of prostate
cancer diagnoses from highly unstructured free-text data.
In conjunction with this resource, it has been possible to
validate a simple yet efective method for the extraction of
relevant information from these types of documents with a
very promising overall performance.</p>
      <p>Annex A contains a complete example of one of the
analysed documents, corresponding to the following textual
fragment from the diagnosis section. The table 1 below
includes the most relevant features extracted from the whole
document. All the characteristics extracted are specified in
the table 4 in the annex.</p>
      <p>BIOPSIA DE PRÓSTATA TRANSPERINEAL, [...]
GRADO DE GLEASON: 7 (3+4) - GRADO GRUPO: 2
[...] PATOLOGÍA ADICIONAL PROSTÁTICA: NO SE
OBSERVA.</p>
      <p>The distribution of cancer presence in the population
studied through pathology reports is depicted in figure 3.
According to the experts, it aligns with the typical average
distribution of PCa diagnoses in the specific region under
study with the characteristics described in the materials
section, remaining about a 2% of the records (370 reports
approximately) to be reviewed by doctors. ASAP, PIN,
inlfammation, and atrophy are all pathologies related to the
possible development of prostate cancer. However, for
diagnostic purposes, medical experts consider them as negative
cases of PCa.</p>
      <p>The distributions of Gleason Score sum and group
retrieved for positive PCa cases are also reproduced in table 2
and figure 4.</p>
      <p>Regarding the compiled corpus of documents, it was
found that the inclusion of the macroscopic section aided in
extracting features related to procedures. The distribution
of procedures found is depicted in figure 5.</p>
      <p>Figure 6 displays the organs directly related to prostate
cancer mentioned in the textual corpus. These mentions are
valuable for characterising and filtering data in downstream
tasks.</p>
      <p>Tumour morphology entities extracted with the NER
model were not found to be helpful at this level of
language analysis, regarding the techniques integrated within
the system. Neither prostate cancer nor Gleason scores
were improved, as these mentions, when correctly detected,
co-occurred with specific prostate domain expressions
already matched in the lexical-morphological analysis. The
transformer-based model was excluded from the process in
order to reduce the system’s runtime.</p>
      <p>Finally, to ensure the system’s correct performance,
prostate cancer diagnosis from pathology reports were
validated against the validation set, i.e. the manually-annotated
morphology dataset, using the predefined tags outlined in
section 3.2: PCa+, PCa- and Untagged. Although, initially,
the morphology coding was not entirely consistent, after
the reviewing cycles it was found to be a useful baseline
and provided valuable support for the detection of complex
cases. These cases were then personally reviewed by
medical experts, thus enhancing our methodology and allowing
for a thorough evaluation of the obtained results.</p>
      <p>As shown in table 3, with respect to the evaluable
pathology reports coinciding in both datasets, a final 96.095%
agreement between NLP-based classification and human
labelling was reached, which corresponds with a weighted
Cohen’s Kappa of 0.9115. During the validation cycles, our
automatic analysis and processing was found to be more
robust than the manual annotation registered in the hospital
repositories, and both were iteratively improved. Gleason
scores, medical procedures, organ mentions and the rest
of the additional markers retrieved also underwent a
cursory validation by human expert intervention, as there was
no corresponding information conveniently categorised in
the clinical repositories or in the previous work studied.
The distribution of data obtained corresponds to the actual
distribution of cases treated in the time period analysed,
as determined by the vast experience of the clinical staf
involved.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and future work</title>
      <p>In this paper we have presented our approach to retrieving,
classifying and structuring information using a combination
of NLP techniques over prostate cancer pathology reports
in Spanish from the records belonging to the Health Sector
Zaragoza II. The methodology designed, and the system
implemented, set the starting point for the development of
a global and homogeneous cancer risk prediction system
based on the biological footprint of patients distributed in
the diferent electronic health records available in the
hospital repositories, supporting and improving the eficiency of
decision support tools for early disease detection processes.</p>
      <p>We have built an extensive structured dataset that serves
to enhance the predictive capacity and explainability of
advanced predictive models for PCa risk identification, such
as multimodal algorithms that work with biological data and
images, and enrich hospital repositories. Additionally, the
validation performed along with the clinical experts throws
very encouraging results, approximating to 96% of annotator
agreement. Nevertheless, a comprehensive and rigorous
validation of the remaining characteristics is still required,
despite their initial alignment with the expectations and
needs of the experts.</p>
      <p>The developed system has been successfully adapted and
executed into a colorectal cancer scenario within this project,
by simply defining an in-domain thesaurus. This has
allowed the processing and evaluation of pathology reports
and then, extending the approach to colonoscopies and
EHR, serving as an efective system to annotate customized
datasets for further research.</p>
      <p>As future work, several avenues are considered, from
the NLP perspective we plan to improve the techniques
explored, delving into higher levels of language analysis as the
semantics, and exploring the new possibilities ofered by the
leading-edge Large Language Models (LLMs). Furthermore,
for the use case under study, the methodology will be
applied to the magnetic resonance reports of prostate cancer
patients as well, with the necessary adaptations for that
specific context, in order to further enhance the dataset
constructed. Lastly, our objective is to extrapolate the analysis
and processing to all types of reports with textual content
to build a comprehensive system that allows for the
identiifcation of a patient’s biological footprint and its influence
on the final diagnosis.</p>
    </sec>
    <sec id="sec-6">
      <title>Ethical Statement</title>
      <p>This study and the use of patient data was approved by the
regional ethics committee of Aragón.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This research was funded by project MIA.2021.M02.0007 of
NextGenerationEU program and Integration and
Development of Big Data and Electrical Systems (IODIDE) group of
Aragon Goverment program.
Diagnosis</p>
      <p>BIOPSIA DE PRÓSTATA TRANSPERINEAL, PROTOCOLO
EXTENDIDO (BASADO EN CAP JUN): -
ADENOCARCINOMA CONVENCIONAL. - AI4: ÁPEX IZQUIERDO,
TRANSICIONAL: - CILINDROS AFECTADOS / REMITIDOS: 1/1
GRADO DE GLEASON: 6 (3+3) - GRADO GRUPO: 1 -
PORCENTAJE DE PATRÓN GLEASON 4 O 5: 0% - PORCENTAJE
DE TEJIDO PROSTÁTICO AFECTADO POR TUMOR: 16,6%
- MM DE CARCINOMA / MM DE CILINDRO: 1/6 MM
AI5: ÁPEX IZQUIERDO, ANTERIOR: - CILINDROS
AFECTADOS / REMITIDOS: 2/2 - GRADO DE GLEASON: 6 (3+3)
GRADO GRUPO: 1 - PORCENTAJE DE PATRÓN GLEASON
4 O 5: 0% - PORCENTAJE DE TEJIDO PROSTÁTICO
AFECTADO POR TUMOR: 61,5% - MM DE CARCINOMA / MM
DE CILINDRO: 8/13 MM - AD4: ÁPEX DERECHO,
TRANSICIONAL: - CILINDROS AFECTADOS / REMITIDOS:2/2
GRADO DE GLEASON: 7 (3+4) - GRADO GRUPO: 2 -
PORCENTAJE DE PATRÓN GLEASON 4 O 5: 5% - PORCENTAJE
DE TEJIDO PROSTÁTICO AFECTADO POR TUMOR: 47,3%
- MM DE CARCINOMA / MM DE CILINDRO: 9/19 MM -
INFILTRACIÓN GRASA PERIPROSTÁTICA: NEGATIVA. -
INFILTRACIÓN DE VESÍCULA SEMINAL: NO VALORABLE
POR AUSENCIA DE VESÍCULA SEMINAL EN EL MATERIAL
REMITIDO. - INVASIÓN LINFOVASCULAR: NEGATIVA.
INVASIÓN PERINEURAL: NEGATIVA. - PATOLOGÍA
ADICIONAL PROSTÁTICA: NO SE OBSERVA.</p>
      <p>Macroscopic findings</p>
      <p>A.- BD1: BASE DERECHO, PERIFÉRICO POSTERIOR: SE
RECIBE UN CILINDRO. INCLUSIÓN TOTAL EN BLOQUE A1.
B.- BD2: BASE DERECHO, PERIFÉRICO EXTERNO: SE RECIBE
UN CILINDRO. INCLUSIÓN TOTAL EN BLOQUE B1. C.- BD3:
BASE DERECHO, PERIFÉRICO ANTERIOR: SE RECIBE UN
CILINDRO. INCLUSIÓN TOTAL EN BLOQUE C1. D.- BD4:
BASE DERECHO, TRANSICIONAL: SE RECIBE UN CILINDRO
FRAGMENTADO. INCLUSIÓN TOTAL EN BLOQUE D1.
E.BD5: BASE DERECHO, ANTERIOR: SE RECIBE UN CILINDRO.
INCLUSIÓN TOTAL EN BLOQUE E1. F.- MD1: MEDIO
DERECHO, PERIFÉRICO POSTERIOR: SE RECIBE UN CILINDRO
MÁS UN FRAGMENTO. INCLUSIÓN TOTAL EN BLOQUE
F1. G.- MD2: MEDIO DERECHO, PERIFÉRICO EXTERNO:
SE RECIBE UN CILINDRO. INCLUSIÓN TOTAL EN BLOQUE
G1. H.- MD3: MEDIO DERECHO, PERIFÉRICO ANTERIOR:
SE RECIBE UN CILINDRO MÁS FRAGMENTO. INCLUSIÓN
TOTAL EN BLOQUE H1. I.- MD4: MEDIO DERECHO,
TRANSICIONAL: SE RECIBE UN CILINDRO. INCLUSIÓN TOTAL
EN BLOQUE I1. J.- MD5: MEDIO DERECHO, ANTERIOR:
SE RECIBE UN CILINDRO. INCLUSIÓN TOTAL EN BLOQUE
J1. K.- AD1: ÁPEX DERECHO, PERIFÉRICO POSTERIOR:
SE RECIBE UN CILINDRO. INCLUSIÓN TOTAL EN BLOQUE
K1. L.- AD2: ÁPEX DERECHO, PERIFÉRICO EXTERNO: SE
RECIBE UN CILINDRO PEQUEÑO. INCLUSIÓN TOTAL EN
BLOQUE L1. M.- AD3: ÁPEX DERECHO, PERIFÉRICO
ANTERIOR: SE RECIBE UN CILINDRO (EN VARIOS
FRAGMENTOS). INCLUSIÓN TOTAL EN BLOQUE M1. N.- AD4: ÁPEX
DERECHO, TRANSICIONAL: SE RECIBE UN CILINDRO.
INCLUSIÓN TOTAL EN BLOQUE N1. O.- AD5: ÁPEX
DERECHO, ANTERIOR: SE RECIBE UN CILINDRO MUY FINO.
INCLUSIÓN TOTAL EN BLOQUE O1.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Rabaan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Bakhrebah</surname>
          </string-name>
          , H. AlSaihati, S. Alhumaid,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Alsubki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Turkistani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Al-Abdulhadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Aldawood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Alsaleh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. N.</given-names>
            <surname>Alhashem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Almatouq</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Alqatari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. E.</given-names>
            <surname>Alahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Sharbini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Alahmadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alsalman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alsayyah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Mutair</surname>
          </string-name>
          ,
          <article-title>Artificial intelligence for clinical diagnosis and treatment of prostate cancer</article-title>
          ,
          <source>Cancers</source>
          <volume>14</volume>
          (
          <year>2022</year>
          )
          <article-title>5595</article-title>
          . URL: http://dx.doi.org/10.3390/cancers14225595. doi:
          <volume>10</volume>
          .3390/cancers14225595.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Chuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Su</surname>
          </string-name>
          , D.-H. Han,
          <string-name>
            <given-names>Y.-W.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-C.</given-names>
            <surname>Lee</surname>
          </string-name>
          , Y.-F. Cheng, T.-P. Hong,
          <string-name>
            <surname>K. S.-M. Li</surname>
            ,
            <given-names>H.-Y.</given-names>
          </string-name>
          <string-name>
            <surname>Ou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
          </string-name>
          , C.-C.
          <article-title>Wang, Efective natural language processing and interpretable machine learning for structuring ct livertumor reports</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>116273</fpage>
          -
          <lpage>116286</lpage>
          . URL: http://dx.doi.org/10.1109/ACCESS.
          <year>2022</year>
          .
          <volume>3218646</volume>
          . doi:
          <volume>10</volume>
          .1109/access.
          <year>2022</year>
          .
          <volume>3218646</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H. R.</given-names>
            <surname>Abdulshaheed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A. Mohammed</given-names>
            <surname>Al-Juboori</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. A</surname>
          </string-name>
          . Al Sayed,
          <string-name>
            <given-names>I. A.</given-names>
            <surname>Barazanchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Gheni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z. A.</given-names>
            <surname>Jaaz</surname>
          </string-name>
          ,
          <article-title>Research on optimization strategy of medical data information security and privacy</article-title>
          , in: 2022 9th International Conference on Electrical Engineering, Computer Science and Informatics (EECSI), IEEE,
          <year>2022</year>
          . URL: http:// dx.doi.org/10.23919/EECSI56542.
          <year>2022</year>
          .
          <volume>9946606</volume>
          . doi:
          <volume>10</volume>
          . 23919/eecsi56542.
          <year>2022</year>
          .
          <volume>9946606</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Houssein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <article-title>Machine learning techniques for biomedical natural language processing: a comprehensive review</article-title>
          ,
          <source>IEEE Access 9</source>
          (
          <year>2021</year>
          )
          <fpage>140628</fpage>
          -
          <lpage>140653</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Natural language processing applications for computer-aided diagnosis in oncology</article-title>
          ,
          <source>Diagnostics</source>
          <volume>13</volume>
          (
          <year>2023</year>
          ). URL: https: //www.mdpi.com/2075-4418/13/2/286. doi:
          <volume>10</volume>
          .3390/ diagnostics13020286.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alawad</surname>
          </string-name>
          , H.
          <article-title>-</article-title>
          <string-name>
            <surname>J. Yoon</surname>
            ,
            <given-names>G. D.</given-names>
          </string-name>
          <string-name>
            <surname>Tourassi</surname>
          </string-name>
          ,
          <article-title>Coarse-to-fine multi-task training of convolutional neural networks for automated information extraction from cancer pathology reports</article-title>
          ,
          <source>in: 2018 IEEE EMBS International Conference on Biomedical &amp;amp; Health Informatics (BHI)</source>
          , IEEE,
          <year>2018</year>
          . URL: http://dx.doi.org/10.1109/BHI.
          <year>2018</year>
          .
          <volume>8333408</volume>
          . doi:
          <volume>10</volume>
          .1109/bhi.
          <year>2018</year>
          .
          <volume>8333408</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Information extraction from pathology reports in a hospital setting</article-title>
          ,
          <source>in: Proceedings of the 20th ACM international conference on Information and knowledge management</source>
          ,
          <source>CIKM '11</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2011</year>
          . URL: http://dx.doi.org/10.1145/2063576.2063846. doi:
          <volume>10</volume>
          .1145/2063576.2063846.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>DiBello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Weinmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. E.</given-names>
            <surname>Richert-Boe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Ritzwoller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Vandeneeden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Jacobsen</surname>
          </string-name>
          ,
          <article-title>Development of an algorithm to identify metastatic prostate cancer in electronic medical records using natural language processing</article-title>
          .,
          <source>Journal of clinical oncology : oficial journal of the American Society of Clinical Oncology 32</source>
          <volume>30</volume>
          _
          <issue>suppl</issue>
          (
          <year>2014</year>
          )
          <article-title>164</article-title>
          . URL: https://api.semanticscholar.org/CorpusID:25873512.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gelfond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Slezak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Porter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Jacobsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. W.</given-names>
            <surname>Chien</surname>
          </string-name>
          ,
          <article-title>Extracting data from electronic medical records: validation of a natural language processing program to assess prostate biopsy results</article-title>
          ,
          <source>World Journal of Urology</source>
          <volume>32</volume>
          (
          <year>2014</year>
          )
          <fpage>99</fpage>
          -
          <lpage>103</lpage>
          . URL: https://api.semanticscholar. org/CorpusID:8917027.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>O.</given-names>
            <surname>Hamzeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rueda</surname>
          </string-name>
          ,
          <article-title>A gene-disease-based machine learning approach to identify prostate cancer biomarkers</article-title>
          ,
          <source>in: Proceedings of the 10th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>633</fpage>
          -
          <lpage>638</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Morote</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Schwartzman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Borque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Esteban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Celma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roche</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. M. de Torres</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Mast</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          <string-name>
            <surname>Semidey</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Regis</surname>
          </string-name>
          , et al.,
          <article-title>Prediction of clinically significant prostate cancer after negative prostate biopsy: The current value of microscopic findings</article-title>
          ,
          <source>in: Urologic Oncology: Seminars and Original Investigations</source>
          , volume
          <volume>39</volume>
          ,
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          ,
          <year>2021</year>
          , pp.
          <fpage>432</fpage>
          -
          <lpage>e11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Khosravi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lysandrou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eljalby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kazemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zisimopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sigaras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brendel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Barnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ricketts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Meleshko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yat</surname>
          </string-name>
          , T. D.
          <string-name>
            <surname>McClure</surname>
            ,
            <given-names>B. D.</given-names>
          </string-name>
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sboner</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Elemento</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chughtai</surname>
            ,
            <given-names>I. Hajirasouliha</given-names>
          </string-name>
          ,
          <article-title>A deep learning approach to diagnostic classification of prostate cancer using pathology-radiology fusion</article-title>
          ,
          <source>Journal of Magnetic Resonance Imaging</source>
          <volume>54</volume>
          (
          <year>2021</year>
          )
          <fpage>462</fpage>
          -
          <lpage>471</lpage>
          . URL: http://dx.doi.org/10. 1002/jmri.27599. doi:
          <volume>10</volume>
          .1002/jmri.27599.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Breischneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zillner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hammon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sonntag</surname>
          </string-name>
          ,
          <article-title>Automatic extraction of breast cancer information from clinical reports</article-title>
          ,
          <source>in: 2017 IEEE 30th International Symposium on Computer-Based Medical Systems (CBMS)</source>
          , IEEE,
          <year>2017</year>
          . URL: http://dx. doi.org/10.1109/CBMS.
          <year>2017</year>
          .
          <volume>138</volume>
          . doi:
          <volume>10</volume>
          .1109/cbms.
          <year>2017</year>
          .
          <volume>138</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>H.-J. Yoon</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gounley</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          <string-name>
            <surname>Young</surname>
          </string-name>
          , G. Tourassi,
          <article-title>Information extraction from cancer pathology reports with graph convolution networks for natural language texts</article-title>
          ,
          <source>in: 2019 IEEE International Conference on Big Data (Big Data)</source>
          , IEEE,
          <year>2019</year>
          . URL: http://dx. doi.org/10.1109/BigData47090.
          <year>2019</year>
          .
          <volume>9006270</volume>
          . doi:
          <volume>10</volume>
          . 1109/bigdata47090.
          <year>2019</year>
          .
          <volume>9006270</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Alsentzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Boag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-H.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Naumann</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>McDermott, Publicly available clinical bert embeddings</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>03323</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ginart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Interian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Upadhaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lupo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Braunstein</surname>
          </string-name>
          ,
          <article-title>Oncobert: Building an interpretable transfer learning bidirectional encoder representations from transformers framework for longitudinal survival prediction of cancer patients</article-title>
          ,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .21203/rs.3. rs-
          <volume>3158152</volume>
          /v1.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Seda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. d. P. P.</given-names>
            <surname>León</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Conde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C. G.</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Sánchez</surname>
          </string-name>
          , G. Rodríguez,
          <string-name>
            <given-names>J. A. P.</given-names>
            <surname>Simón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. L. P.</given-names>
            <surname>Calderón</surname>
          </string-name>
          ,
          <article-title>Plataforma para la extracción automática y codificación de conceptos dentro del ámbito de la oncohematología (proyecto coco)</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>61</volume>
          (
          <year>2018</year>
          )
          <fpage>65</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Miranda-Escalada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Farré</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <article-title>Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results</article-title>
          .,
          <source>IberLEF@ SEPLN</source>
          (
          <year>2020</year>
          )
          <fpage>303</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>O.</given-names>
            <surname>Solarte-Pabón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Montenegro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García-Barragán</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Torrente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Provencio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Menasalvas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Robles</surname>
          </string-name>
          ,
          <article-title>Transformers for extracting breast cancer information from spanish clinical narratives</article-title>
          ,
          <source>Artificial Intelligence in Medicine</source>
          <volume>143</volume>
          (
          <year>2023</year>
          )
          <article-title>102625</article-title>
          . URL: https://www.sciencedirect.com/science/article/ pii/S0933365723001392. doi:https://doi.org/10. 1016/j.artmed.
          <year>2023</year>
          .
          <volume>102625</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>M. M. Van Berkum</surname>
          </string-name>
          ,
          <article-title>Snomed ct® encoded cancer protocols</article-title>
          ,
          <source>in: Amia Annual Symposium Proceedings</source>
          , volume
          <volume>2003</volume>
          , American Medical Informatics Association,
          <year>2003</year>
          , p.
          <fpage>1039</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Organization</surname>
          </string-name>
          , Icd-10 :
          <article-title>international statistical classification of diseases and related health problems : tenth revision</article-title>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Carrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Llop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gutiérrez-Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Armengol-Estapé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silveira-Ocampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Valencia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gonzalez-Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <article-title>Pretrained biomedical language models for clinical NLP in Spanish</article-title>
          ,
          <source>in: Proceedings of the 21st Workshop on Biomedical Language Processing</source>
          , Association for Computational Linguistics, Dublin, Ireland,
          <year>2022</year>
          , pp.
          <fpage>193</fpage>
          -
          <lpage>199</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .bionlp-
          <volume>1</volume>
          .19. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2022</year>
          .bionlp-
          <volume>1</volume>
          .
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J. I.</given-names>
            <surname>Epstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. C.</given-names>
            <surname>Allsbrook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Amin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Egevad</surname>
          </string-name>
          ,
          <article-title>The 2005 international society of urological pathology (isup) consensus conference on gleason grading of prostatic carcinoma</article-title>
          ,
          <source>American Journal of Surgical Pathology</source>
          <volume>29</volume>
          (
          <year>2005</year>
          )
          <fpage>1228</fpage>
          -
          <lpage>1242</lpage>
          . URL: http://dx. doi.org/10.1097/01.pas.
          <volume>0000173646</volume>
          .99337.b1. doi:
          <volume>10</volume>
          . 1097/01.pas.
          <volume>0000173646</volume>
          .99337.
          <year>b1</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R. D.</given-names>
            <surname>Rosen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sapra</surname>
          </string-name>
          ,
          <article-title>Tnm classification</article-title>
          .,
          <year>2023</year>
          . URL: https://www.ncbi.nlm.nih.gov/books/NBK553187/, last accessed:
          <fpage>2023</fpage>
          -09-30.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Annex</surname>
          </string-name>
          <article-title>1: Full example</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>