<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Connecting Science Data Using Semantics and Information Extraction</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author role="corresp">
							<persName><forename type="first">Evan</forename><forename type="middle">W</forename><surname>Patton</surname></persName>
							<email>pattoe@cs.rpi.edu</email>
							<affiliation key="aff0">
								<orgName type="institution">Rensselaer Polytechnic Institute</orgName>
								<address>
									<addrLine>110 8 th Street</addrLine>
									<postCode>12180</postCode>
									<settlement>Troy</settlement>
									<region>NY</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Deborah</forename><forename type="middle">L</forename><surname>Mcguinness</surname></persName>
							<affiliation key="aff0">
								<orgName type="institution">Rensselaer Polytechnic Institute</orgName>
								<address>
									<addrLine>110 8 th Street</addrLine>
									<postCode>12180</postCode>
									<settlement>Troy</settlement>
									<region>NY</region>
									<country key="US">USA</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Connecting Science Data Using Semantics and Information Extraction</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">1B456388C2E5D3AB6C1477A25C370950</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T12:52+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>knowledge representation</term>
					<term>explanation</term>
					<term>clinical notes</term>
					<term>natural language</term>
					<term>web forums</term>
					<term>nanopublications</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>We are developing prototypes that explicate our vision of connecting personal medical data to scientific literature as well as to emerging grey literature (e.g., community forums) to help people find and understand information relevant to complex medical journeys. We focus on robust combinations of natural language processing along with linked data and knowledge representation to build knowledge graphs that help people make sense of current conditions and enable new manners of scientific hypothesis generation. We present our work in the context of a breast cancer use case. We discuss the benefits of biomedical linked data resources and describe some potential assistive technology for navigating rich, diverse medical content.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>As scientific knowledge continues to grow in size and diversity, it is increasingly difficult to discover and manage information relevant to any particular context. It can be challenging to determine how a statement or report relates to others and to form and evaluate (often competing) hypotheses, e.g. related to diagnosis or treatment paths. Complications grow when content is both structured and unstructured, and when some is from less accredited sources. We aim to expand the boundaries of Linked Science by focusing on evidence modeling from natural language processing techniques (NLP) over broad content and by identifying promising data-driven hypotheses using linked data and nanopublication style encodings. We present this discussion in the context of a breast cancer demonstration use case informed by challenges experienced during a co-author's recent cancer journey. Cancer is a complex disease to manage and treat, often requiring chemotherapy, surgery, radiation, and drugs to reduce recurrence. We show how management of this information by the patient is aided by semantic technologies combined with natural language processing algorithms.</p><p>A breast cancer patient wishes to better understand her diagnosis and planned treatment. She is interested in expected chemotherapy side effects, and leveraging experiences of other similar individuals to proactively find and evaluate promising coping strategies. She reads through oncologist-provided documents about her proposed chemotherapy drugs and uses search engines to find more about likely adverse effects that appear detrimental to her quality of life. She finds conflicting opinions on the efficacy of different coping strategies, and needs to determine an approach to effectively weigh the possible pros and cons. Managing this information is mentally taxing and can easily overwhelm a patient.</p><p>Our patient needs to find and comprehend potentially conflicting evidence about treatment options and side effects. We propose new software, using a variety of artificial intelligence tools built on the interoperability principles promulgated by linked data and the Semantic Web, to address these challenges.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Evidence Modeling</head><p>The patient uses current technologies to obtain information about her treatment strategy and to formulate promising side effect mitigations. This can be time consuming for anyone, but more so for medically naïve patients. Furthermore, technologies such as web forums or social networking sites are becoming increasingly common for discourse between patients as they can often include anecdotal reports, that have not yet been validated through clinical trials, but may be valuable. They are often presented in layperson terms and sometimes attract new patients who may be less medically literate. Due to lack of scientific rigor, there may be contradictory or unsupported information available, as shown in the following two answers about a mitigation for the very common, taxol-related, nail bed problem:</p><p>My onc[ology] nurse told me to rub tea tree oil into my cuticles and nails every night. It is a natural anti-septic and for whatever reason can sometimes help prevent nail infections and lifting during taxol. <ref type="foot" target="#foot_0">1</ref>I wouldn't use tea tree oil. A friend did on some cracked skin and it got worse. <ref type="foot" target="#foot_1">2</ref>The first suggestion is a common preventive approach for nail problems: tea tree oil prevents nail infections because "it is a natural anti-septic" and appeals to authority "my onc nurse told me to...". The second suggestion from a different user in the same thread advises against tea tree oil as "a friend [applied tea tree oil] on some cracked skin and it got worse." Natural Language techniques may be used to extract coping strategies for particular conditions but without deeper knowledge, provenance, and tools, the user may not know how to evaluate and/or integrate potentially contradictory suggestions. We are extending joint extraction techiques proposed in <ref type="bibr" target="#b3">[4]</ref> with semantic background knowledge to aid in extracting linked data from medical records.</p><p>The Repurposing Drugs using Semantics (ReDrugS) project <ref type="bibr" target="#b4">[5]</ref> has focused on modeling evidence using small units of publishable information called Nanopublications <ref type="bibr" target="#b1">[2]</ref>. ReDrugS utilizes linked data sources to build a knowledge base of nanopublications that is then reasoned about using probabilistic techniques to identify potential links between proteins, drugs, binding sites, and genes, with the ultimate aim of discovering possible new off-label uses for FDA-approved drugs. This project's success has been partially due to the large corpus of linked data and ontologies generated by the biomedical community over the past few decades. ReDrugS has ingested content from 17 structured curated data sources, including content concerning drugs, alternate names, conditions, and pathways. Once a chemotherapy protocol is extracted from medical notes, ReDrugs can be used to find alternative drug names along with related conditions. This framework, along with the side effect resource SIDER in process, can be used to improve the patient's process in finding chemotherapy drug side effects and some mitigations by applying its search techniques to authoritative drug resources, such as looking for anti-nausea prescription drugs. The infrastructure for this system could be repurposed for other scientific domains, but only if linked data sources are abundant in those domains or if quality linked data can be generated from automated methods, e.g. via natural language processing of web-based resources.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Explanations</head><p>We aim to provide extensive explanation mechanisms since explanation is a key component of transparent systems and user studies have shown that explanations are required if agents are to be trusted <ref type="bibr" target="#b0">[1]</ref>. We aid explanation generation through the collection of provenance, modeled using the W3C's PROV ontology <ref type="bibr" target="#b2">[3]</ref>. PROV-O is a standard for modeling provenance information on the web, which allows tools to integrate distributed provenance information from different systems. We use this provenance to help construct end user explanations that include both lineage of content and support (and opposition) for a statement.</p><p>We identify potential evidence on the use of tea tree oil in chemotherapyinduced nail bed problems. Not only would a patient want to know evidence, source, and authoritativeness for both views, she might also want the system further decompose these arguments and present supporting evidence as to the antimicrobial nature of tea tree oil in more authoritative sources (e.g. <ref type="bibr" target="#b5">[6]</ref>).</p><p>We claim that we can reuse the ReDrugS content to find prescription drugs for chemotherapy side effects. Provenance may be displayed to show that the recommendation is from a validated authoritative source. While that framework was originally designed to find potential new off-label uses for drugs along with confidence ratings, the explanation component is more critical for our use so that researchers may inspect evidence sources and the methods used to determine the system confidence. Without such explanations, people would have difficulty evaluating competing suggestions.</p><p>Our systems<ref type="foot" target="#foot_2">3</ref> provide explanation drill down so users can obtain as much detail as they desire, thus allowing a patient to find, for example, if authoritative sources contain prescription drugs for coping with a particular side effect. Our NL-based extraction work can be used to identify alternative, possibly competing, therapies, e.g. an herbal remedy recommended anecdotally with potentially corroborating authoritative sources.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">Discussion and Summary</head><p>Natural Language Processing can expose some of the unstructured content of medical records as structured content as well as assist in generating linked data from unstructured sources. The ReDrugS framework provides a semanticallyintegrated system combining many different structured biomedical resources to generate a broadly reusable knowledge graph. By integrating the natural language and structured knowledge representation approaches, we can obtain a much richer annotated knowledge base that includes source and confidence information. Our prototypes demonstrate some ways that this rich resource may then be used to help patients and their support networks to discover, integrate, and evaluate information relevant to complicated medical situations and to help form transparent and data-driven hypotheses about how to proceed. We believe these efforts demonstrate some opportunities for future AI-enhanced Linked Sciencebased assistants that use the wealth of structured content as well as the growing grey literature collection.</p></div>			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0">https://community.breastcancer.org/forum/69/topic/783573</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1">https://community.breastcancer.org/forum/96/topic/745475</note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_2">http://tw.rpi.edu/web/project/MobileHealth http://tw.rpi.edu/web/project/ReDrugS</note>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Acknowledgements</head><p>The authors thank Heng Ji and Alex Borgida for their discussions that helped shape this work.</p></div>
			</div>

			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">Toward establishing trust in adaptive agents</title>
		<author>
			<persName><forename type="first">A</forename><surname>Glass</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">L</forename><surname>Mcguinness</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Wolverton</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">13th Intl Conference on Intelligent User Interfaces</title>
				<imprint>
			<date type="published" when="2008">2008</date>
			<biblScope unit="page" from="227" to="236" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">The anatomy of a nanopublication</title>
		<author>
			<persName><forename type="first">P</forename><surname>Groth</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Gibson</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Velterop</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Information Services &amp; Use</title>
		<imprint>
			<biblScope unit="volume">30</biblScope>
			<biblScope unit="page" from="51" to="56" />
			<date type="published" when="2010">2010</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<monogr>
		<title level="m" type="main">PROV-O: The PROV ontology</title>
		<author>
			<persName><forename type="first">T</forename><surname>Lebo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Sahoo</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">L</forename><surname>Mcguinness</surname></persName>
		</author>
		<imprint>
			<date type="published" when="2013">2013</date>
			<pubPlace>W3C</pubPlace>
		</imprint>
	</monogr>
	<note type="report_type">Tech. rep</note>
</biblStruct>

<biblStruct xml:id="b3">
	<analytic>
		<title level="a" type="main">Incremental joint extraction of entity mentions and relations</title>
		<author>
			<persName><forename type="first">Q</forename><surname>Li</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Ji</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">Proc. of the 52nd Annual Meeting of the Association for Computational Linguistics</title>
				<meeting>of the 52nd Annual Meeting of the Association for Computational Linguistics</meeting>
		<imprint>
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">A nanopublication framework for systems biology and drug repurposing</title>
		<author>
			<persName><forename type="first">J</forename><surname>Mccusker</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Solanki</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Chang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Dumontier</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Dordick</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">L</forename><surname>Mcguinness</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">CSHALS</title>
		<imprint>
			<biblScope unit="volume">2014</biblScope>
			<date type="published" when="2014">2014</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">A review of applications of tea tree oil in dermatology</title>
		<author>
			<persName><forename type="first">N</forename><surname>Pazyar</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Yaghoobi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">N</forename><surname>Bagherani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Kaerouni</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">International Journal of Dermatology</title>
		<imprint>
			<biblScope unit="page" from="784" to="790" />
			<date type="published" when="2013">2013</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
