Environmental Health Language Collaborative (EHLC): a route to environmental health science data harmonization Anna Maria Masci 1, Stephanie Holmgren 1, Charles Schmitt1, Rima Habre2, Anne E Thessen3, Rebecca Boyles4, Carmen Marsit5. 1 Office of Data Science, National Institute of Environmental Health Sciences (NIEHS), Research Triangle Park, North Carolina, USA 2 University of Southern California, Los Angeles, CA, USA 3 University of Colorado Anschutz Medical Campus, Center for Health AI, Aurora CO, USA 4 Center for Data Modernization Solutions, RTI International. Durham, NC. USA 5 Gangarose Department of Environmental Health, Emory University Rollins School of Public Health, Atlanta, GA, USA Abstract Standard language is critical for helping scientists share, compare, and reanalyze data. The increased use of automatization and AI technologies has made the adoption of machine interpretable language essential. Due to the broadness of the domains that are under the environmental health umbrella there is not yet a set of common standard terminologies. To address this lack of standardized language, NIEHS has launched the Environmental Health Language Collaborative (EHLC) https://www.niehs.nih.gov/research/progra ms/ehlc/index.cfm. This is a new initiative to advance community development and application of a harmonized language for describing Environmental Health Science (EHS) research. As a first step toward the development of standard terminology, a working group of environmental health researchers and NIEHS program officers established an initial set of four general use cases. Here we present one of the initial use cases on place- based exposures. This preliminary work is intended to be expanded as the community develops. EHLC is seeking larger community involvement as well as additional use cases. NIEHS encourages anyone interested in advancing this mission to engage in this community. Keywords Ontology; controlled vocabulary; data reuse; FAIR data metadata; taxonomy; standards; semantic; environmental health; toxicology; community of practice; community driven, geospatial, place, location, exposure. CEUR ceur-ws.org Workshop ISSN 1613-0073 Proceedings 1 1. Introduction are inconsistencies and gaps in the terminologies and ontologies used within Environmental health (EH) is a science that subfields, but scientific language is often studies the effect of exposure to domain-specific and standardizing or even environmental factors on human health. The harmonizing language across subfields is definition of environment is wide and especially challenging. includes the “totality of exposures we face To address the need for common language, throughout our lives, e.g., the food we ingest, NIEHS has launched the Environmental the air we breathe, the objects we touch, the Health Language Collaborative (EHLC)[4] psychological stresses we face, the activities https://www.niehs.nih.gov/research/program in which we engage” [1] EH research is not s/ehlc/index.cfm. This is a new initiative to just focused on external exposures, but also advance community development and considers the molecules in our body that application of a harmonized language for derive from external exposures, the describing Environmental Health Science environmental influences we receive through (EHS) research. our parents, the socio-economic factors that play into disparities in health as well as The proposed mission of this community is research that seeks to remediate and reduce to: the impact of these factors, e.g., by engineering plants that can remove or reduce • Apply language standards and best pollutants. practices for accurate environmental The EH field covers a diversity of domains health data and knowledge and methodologies, such as environmental representation epidemiology, toxicology, clinical and • Cultivate a vocabulary aware translational research, immunology, environmental health community microbiology, exposure science, social through training and education science, and environmental engineering. • Foster community-based Progress in EH research depends on the development of harmonized ability to compare, contrast, and integrate data vocabularies, terminologies, and from across the field, which requires adoption ontologies of the principles of Findable, Accessible, • Identify use cases for applying Integrable, and Reusable (FAIR) [2, 3]data. knowledge organization systems in FAIR requires the use of either common or research comparable language in describing scientific • Promote and develop methods and data, metadata, and findings. tools for applying harmonized The breadth of the EH field challenges the use language in research of a common language, not only because there ICBO 2022, September 25–28, 2022, Ann Arbor, MI, USA EMAIL: mascia2@niehs.nih.gov (Anna Maria masci.); holmgre1@niehs.nih.gov (Stephanie Holmgren.); charles.schmitt@nih.gov (Charles Schmitt.); habre@usc.edu ( Rima Habre.); annethessen@gmail.com (Anne E. Thessen); rboyles@rti.org ( Rebecca Boyles); carmen.j.marsit@emory.edu (Carmen Marsit). ORCID: 0000-0003-1940-6740 (Anna Maria Masci.); 0000- 0002-3148-2263 (Charles Schmitt); 0000-0002-2908-3327 ( Anne E. Thessen); 0000-0003-0073-6854 (Rebecca Boyles); 0000- 0003-4566-150X (Carmen Marsit). 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). CEUR Workshop Proceedings (CEUR-WS.org) ©️ A first step towards a practical approach has 5. What do my unique exposure been asking what scientific questions would conditions based on where I live and benefit most from development and adoption work (E.g., Geographical Location, of a harmonized language standard? A Occupation, Regulations, Hobbies) working group of environmental health indicate about potential risks to my researchers and NIEHS program officers health? developed an initial set of five general use Due to space limitation, we present the very cases examples as starting points for preliminary work done for use case five to community discussion. Working groups led begin building the ontological representation by use case champions were formed from the of environmental exposures and social community to address each use case. The stressors or factors assessed based on place working group teams decided to focus all the and geospatial information. use cases on the effect of Particulate Matter (PM) main component of the air pollutants on 2. Methods and Results Asthma as a common theme to unite their work. ‘Asthma is a disease of the respiratory Geospatial data are composed of three general tract which is caused by a combination of components Object, Event, Location, and environmental and genetic factors’ [5, 6]. each of these components has specific ‘Particulate matter is an environmental material characteristics that are time related Fig (1). which is composed of microscopic portions of solid or liquid material suspended in another environmental material’[7, 8]. PM derives from multiple different sources, such as vehicle and industrial emissions from fossil fuel combustion, cigarette smoke, and burning organic matter, such as wildfires, as well as chemical reactions that can form PM from precursors. WHO has estimated that 4.2 million deaths occur as a result of exposure to ambient (outdoor) air pollution. ( https://www.who.int/health-topics/air- pollution - tab=tab_2 ). Figure 1. The main components associated Although all the five use cases focus on with geospatial data. Asthma and PM, each of them is trying to answer different questions: Members of the Geospatial Working Group 1. What data exists for a given started by looking at an available data set chemical/endpoint/exposure from the Personalized Environment and scenario? Genes Study (PEGS) 2. How best to combine data from (https://www.niehs.nih.gov/research/clinical/ multiple independent studies? studies/pegs/index.cfm). 3. Given measures of biological The study’s panel of experts had already responses to one or more exposures, identified the initial set of essential what are the biological processes that components to represent the geospatial data. might be related to the observed changes? Because we are using an existing list of data 4. What are the biomarkers, phenotypes, elements, the first step was to look at the OBO and/or outcomes that can be measured Foundry ontologies [9] to see if those data and used as an indicator of exposure? elements were already captured in existing ontologies. Figure 2 shows a representation of the data elements that were captured and their We then explored the ability of ontology to relations. The different box color represents capture more specific geographical types of the different ontologies from which the terms information. Figure 3 shows a list of terms were imported: Exposure Ontology (ExO) related to geographic location. There are [10], Gazetteer (GZ) terms like latitude measurement datum, (http://environmentontology.github.io/gaz/), longitude measurement datum that have been Ontology of Biomedical Investigations (OBI) already described in Ontology of Biomedical [11, 12], phenotype and trait ontology Investigations. Other terms like geographical (PATO) [13]. In italic are the relations from identifier (GEO ID) and Buffer zone, which the relation Ontology (RO) [14] that we have are commonly used in geospatial studies, are used to link the terms. In addition to the not present in any ontology. imported terms, new terms have been The Census Bureau and other state and identified as ‘stressor detection assay’ and federal agencies are responsible for assigning ‘stressor detector’. For the stressor detection geographic identifiers, or GEOIDs, to assay we are proposing the following geographic entities to facilitate the definition: ‘an assay that aims to detect organization, presentation, and exchange of exposure stressor’. We are proposing the geographic and statistical data. stressor detection to be a child of a more (https://www.census.gov/programs- general term assay defined in OBI. surveys/geography/guidance/geo- identifiers.html) We have classified the GEO ID term as identifier class defined in the IAO (https://obofoundry.org/ontology/iao.html). We have modified the Census Bureau definition for the GEO ID to be ‘is an identifier composed by numeric codes that uniquely identify all administrative/legal and statistical geographic areas for which the Figure 2. Ontological representation of an Census Bureau tabulates data.’ exposure event. The different box colors Another term that we needed to represent is a indicate the different ontologies from which Buffer Zone. This is a very common term used the terms were imported. The gray boxes to define a zone and its characteristics, that indicate the term has not been found in any are the object of the study. Although there is ontology. The green filled box highlights the this term in ENVO its classification under term that is present in Figure 2 as well as administrative region does not fit with our Figure 3. usage of the term. In our use case the buffer zone is used to define a zone from which The second additional term is a stressor collecting data (point, line, area) that is detector. equidistant from the stressor. We are proposing the following definition: Is As is shown in Figure 3 classification of this a role that inheres in a material entity, and term is still under discussion as well as how which is realized through a process of to relate it to a specific geographic location. exposure stressor detection. These two new terms as well as their definitions have been proposed to the ontology community. The red triangle in Figure 2 represents the term that is the linking node between Figures 2 and 3. Environmental Health Language Collaborative. Members of Environmental Health Language Collaborative Geospatial working group 5. References Figure 3. Ontological representation of the geographical specific entities. The different Uncategorized References box colors indicate the different ontologies from which the terms were imported. The 1. Miller, G., The Exposome: A primer gray boxes indicate terms that have not been 2014, https://doi.org/10.1016/C2013- found in any ontology. The green filled box 0-06870-3: Elseview Inc. highlights the term that links Figure 3 and 2. Wilkinson, M.D., et al., The FAIR Figure 2. Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. 3. Discussion 3. Wilkinson, M.D., et al., Addendum: The FAIR Guiding Principles for scientific data management and Geospatial studies use several heterogeneous stewardship. Sci Data, 2019. 6(1): p. data types. Although some efforts have been 6. developed to create a common language like 4. Holmgren, S.D., et al., Catalyzing the Open Geospatial Consortium (OGC) Knowledge-Driven Discovery in (https://www.ogc.org/), there is still a lack of Environmental Health Sciences standardized machine-readable terminology. through a Community-Driven As ontology good practice we are reusing Harmonized Language. Int J Environ several terms from other OBO Foundry Res Public Health, 2021. 18(17). ontologies. We are creating new relations 5. Schriml, L.M., et al., The Human between these terms to better represent Disease Ontology 2022 update. geospatial information. For the terms that Nucleic Acids Res, 2022. 50(D1): p. were not found in other ontologies we have D1255-D1261. created a new term and definitions. 6. Schriml, L.M., et al., Human Disease This is a preliminary attempt to use an Ontology 2018 update: classification, ontological representation of the geospatial content and workflow expansion. data for an exposure. Our near future goal for the use cases is to extract and represent the Nucleic Acids Res, 2019. 47(D1): p. minimal information necessary for capturing use D955-D962. cases data. That could serve as reference for the 7. Buttigieg, P.L., et al., The environmental data environment ontology: These efforts are being developed under the contextualising biological and community-driven Environmental Health biomedical entities. J Biomed Language Collaborative to ensure an open Semantics, 2013. 4(1): p. 43. and broad community participation and 8. Buttigieg, P.L., et al., The development of harmonized language environment ontology in 2016: bridging domains with increased 4. Acknowledgements scope, semantic density, and interoperation. J Biomed Semantics, 2016. 7(1): p. 57. 9. Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. 10. Mattingly, C.J., et al., Providing the missing link: the exposure science ontology ExO. Environ Sci Technol, 2012. 46(6): p. 3046-53. 11. Bandrowski, A., et al., The Ontology for Biomedical Investigations. PLoS One, 2016. 11(4): p. e0154556. 12. Vita, R., et al., Standardization of assay representation in the Ontology for Biomedical Investigations. Database (Oxford), 2021. 2021. 13. Gkoutos, G.V., P.N. Schofield, and R. Hoehndorf, The anatomy of phenotype ontologies: principles, properties and applications. Brief Bioinform, 2018. 19(5): p. 1008- 1021. 14. Smith, B., et al., Relations in biomedical ontologies. Genome Biol, 2005. 6(5): p. R46.