<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Semantic Management of Data from Biodiversity and Ecosystem Studies: Toward an Integrated Workflow from Collection to Publication. Application to Plankton Data from Lake Geneva</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Christian</forename><surname>Pichot</surname></persName>
							<email>christian.pichot@inrae.fr</email>
							<affiliation key="aff0">
								<orgName type="laboratory">INRAE</orgName>
								<orgName type="institution">URFM</orgName>
								<address>
									<addrLine>228 route de l&apos;Aérodrome</addrLine>
									<postCode>84914</postCode>
									<settlement>Avignon</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Damien</forename><surname>Maurice</surname></persName>
							<email>damien.maurice@inrae.fr</email>
							<affiliation key="aff1">
								<orgName type="laboratory">INRAE</orgName>
								<orgName type="institution">UMR SILVA</orgName>
								<address>
									<addrLine>route d&apos;Amance</addrLine>
									<postCode>54280</postCode>
									<settlement>Champenoux</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Ghislaine</forename><surname>Monet</surname></persName>
							<email>ghislaine.monet@inrae.fr</email>
							<affiliation key="aff2">
								<orgName type="laboratory" key="lab1">INRAE</orgName>
								<orgName type="laboratory" key="lab2">UMR CARRTEL</orgName>
								<address>
									<addrLine>75 avenue de Corzent</addrLine>
									<postCode>74200</postCode>
									<settlement>Thonon-les-bains</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Rachid</forename><surname>Yahiaoui</surname></persName>
							<email>rachid.yahiaoui@inrae.fr</email>
							<affiliation key="aff3">
								<orgName type="laboratory">INRAE</orgName>
								<orgName type="institution">US INFOSOL</orgName>
								<address>
									<addrLine>2163 avenue de la Pomme de Pin</addrLine>
									<postCode>45075</postCode>
									<settlement>Orléans</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Philippe</forename><surname>Clastre</surname></persName>
							<email>philippe.clastre@inrae.fr</email>
							<affiliation key="aff0">
								<orgName type="laboratory">INRAE</orgName>
								<orgName type="institution">URFM</orgName>
								<address>
									<addrLine>228 route de l&apos;Aérodrome</addrLine>
									<postCode>84914</postCode>
									<settlement>Avignon</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Benjamin</forename><surname>Jaillet</surname></persName>
							<email>benjamin.jaillet@inrae.fr</email>
							<affiliation key="aff0">
								<orgName type="laboratory">INRAE</orgName>
								<orgName type="institution">URFM</orgName>
								<address>
									<addrLine>228 route de l&apos;Aérodrome</addrLine>
									<postCode>84914</postCode>
									<settlement>Avignon</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Semantic Management of Data from Biodiversity and Ecosystem Studies: Toward an Integrated Workflow from Collection to Publication. Application to Plankton Data from Lake Geneva</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">6D2ED7248542D1F6AA3ADA4D0B2ED0B3</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-24T12:08+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>interoperability</term>
					<term>biodiversity</term>
					<term>plankton</term>
					<term>ontology</term>
					<term>modelling</term>
					<term>pipeline</term>
					<term>entity property</term>
					<term>FAIR data</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Biodiversity is a key player in ecosystem characteristics and dynamics. Acting as a driver, it also results from ecosystem functioning. Understanding this complex interplay between biological and physical components is one of the main current challenges in the context of land use changes and climate warming. The acquisition of knowledge on biodiversity requires multidisciplinary approaches and mobilises numerous research teams. Data are collected or computed in large quantity but are most often poorly standardised and therefore heterogeneous. In this context the development of semantic interoperability is a major challenge for the sharing and reuse of these data. This objective is implemented within the framework of the AnaEE (Analysis and Experimentation on Ecosystems) Research Infrastructure dedicated to experimentation on ecosystems and biodiversity. A distributed Information System (IS) is developed, based on the semantic interoperability of its components using common vocabularies (AnaeeThes thesaurus and OBOE-based ontology extended for disciplinary needs) for modelling the studied system. This modelling covers the measured variables including biodiversity, as well as the different components of the experimental or observational context, from sensor to plot and network. Driven by the ontology, the approach relies on the atomic decomposition of each of the components into observed entities, their characteristics and qualifiers, their units or naming standards. The modelling of the system allows the semantic annotation of relational databases or flat files for the production of URIs based graph databases. A first pipeline automates the annotation process and the production of the semantic data. A second pipeline is devoted to the exploitation of these semantic data by generating i) metadata records formatted according to the geospatial extension for the Data Catalog Vocabulary standard and the ISO 19139 standard, and ii) Network Common Data Form data files. The implementation of this integrated semantic management of data is presented here for phytoand zoo-plankton data collected from water columns in Lake Geneva over a 30 years period, as well as for environmental data about water temperature and nutrients. The work carried out contributes to the development and use of semantic vocabularies within the biodiversity and ecology research community, leading to semantically enriched metadata records and interoperable data sets. The genericity of the tools make them usable in different contexts of data production, management and ontologies involved in semantic modelling.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>The knowledge of ecosystem structure and functioning is more than never a prerequisite to tackle the global challenges we are now facing (global warming, food supply, biodiversity preservation...). Actually we have to anticipate the middle to long term trajectory of our living planet according to the socio-economic-political choices that could be made. Lessons from the past and empirical knowledge are of course to be taken into account but the complexity of the system in which multiple interactions take place and, above all, the unprecedented environmental context produced by global warming, make it essential to increase scientific knowledge and share it across disciplines. Acting as a driver as well as resulting from ecosystem functioning biodiversity takes a central place in the game. In addition to data collection, their FAIRification for common understanding, sharing and re-use is the challenge.</p><p>In this general context, the AnaEE Research Infrastructure develops services dedicated to the study of continental terrestrial and aquatic ecosystems. Its thematic scope concerns the biological diversity and functioning of grassland, crop, forest and lake ecosystems. The services offered include open-air and closed experimental platforms, analytical platforms and digital resources for data management and data-model coupling <ref type="bibr" target="#b0">[1]</ref>. According to the studied ecosystem, various qualitative or quantitative features of biodiversity are analysed such as: flora and resident soil organisms in grassland; soil microbial diversity; phytoplankton or fishes in freshwater ecosystems.</p><p>Produced by distributed platforms, most of the collected data are initially poorly standardised and managed in different information systems from flat files to relational databases. With the objective of ensuring technical and semantic interoperability a workflow is developed for data annotation and exploitation. We present here the general strategy deployed and provide examples of its implementation for biodiversity (zooplankton and phytoplankton) and environment data (water temperature and phosphorus concentration) collected from Lake Geneva, at a single sampling point referred to as SHL2, the deepest and pelagic part of the lake <ref type="bibr" target="#b1">[2]</ref>. Data were collected once a month in winter and twice a month in other seasons.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Strategy and Workflow</head><p>The publication of open access data (FAIR) is nowadays easy. It is also becoming easier to make people aware of their existence (FAIR) thanks to the catalogues of the data repositories and the interoperability of the metadata standards they used. Nevertheless, their fine description, specific to the relevant thematic field and, even more so, their semantic interoperability (FAIR) are often weaker because requiring a significant investment and the use of shared vocabularies. To be reusable the data set must be described not only for the variables it contents but also for all the surrounding factors that influence variables values. The investment required is all the more important as this work is carried out late, i.e. at data publication step and not during the data life cycle, from acquisition, curation, processing...This analysis motivates the implementation of a workflow operating as far upstream as possible and whose genericity allows it to be used widely.</p><p>The strategy developed (Figure <ref type="figure" target="#fig_0">1</ref>) is based on the modelling of the whole system (the experimental design in AnaEE) using an ontology as described in 3.1. The modeled system is used by a first pipeline <ref type="bibr" target="#b2">[3]</ref> that automates the annotation process of relational database or flat files and generates the rdf triples (data lifting). A second pipeline is devoted to the exploitation of these semantic data by generating i) metadata records formatted according to the geospatial extension for the Data Catalog Vocabulary (GeoDCAT) standard and the ISO 19139 standard, and ii) Network Common Data Form (NetCDF) data files. It also offers a data publication service presently implemented for Dataverse repositories. The content of the dataset is determined by the user of 'pipeline 2' according to criteria defining the perimeter of interest and currently based on variables and variable categories, years, experimental platforms and networks, ecosystems. In addition to the general information on user and date, the defined perimeter is also used for the feeding of the metadata fields of the generated GeoDCAT record (abstract as dcterms:description, keyword and thesaurus as dcat:theme, dcat:contactPoint, dcterms:spatial and dcterms:temporal). </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Ontology Based Modelling</head><p>The choice of the reference ontology was determined by two main objectives: i) the need to model the whole observation or experimentation system and ii) the wish to achieve an atomic of "entityquality" type. As a consequence we adopted the Extensible Observation Ontology (OBOE) <ref type="bibr" target="#b3">[4]</ref> a formal ontology for capturing the semantics of scientific observation and measurement and developed in the frame of ecology <ref type="bibr" target="#b4">[5]</ref>. In addition to the modelling of the observed variable (e.g: water temperature) OBOE can characterize, using the hasContext property, the context of an observation (e.g., space and time), as well as dependencies such as nested experimental observations. The main concepts in OBOE (Figure <ref type="figure" target="#fig_1">2</ref>) include: -Observation: an event in which one or more measurements are taken -Measurement: the measured value of a property for a specific object or phenomenon (e.g., 12.5) -Entity: an object or phenomenon on which measurements are made (e.g., water) -Characteristic: the property being measured (e.g., Temperature -Standard: units and controlled vocabularies for interpreting measured values (e.g., degree celcius) -Protocol: the procedures followed to obtain measurements -Qualifier: statistical process ( e.g., maximum; half-hourly-average) .</p><p>Our modelling covers the measured variables, the different components of the experimental context, from the sensor to the plot and the network, through the atomic decomposition of the observed entities, their characteristics and qualifiers, the units and the naming standards. In order to cover the whole system, OBOE extensions were realised mostly on experimental entity (network, site, plot, treatment), characteristics and standard naming for the experimental components (e.g. site names) and the observable properties (presently 350 variables, e.g dissolved orthophosphorus mass concentration).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Modelling of the Observable Property</head><p>Measured variables are often named succinctly in databases, in a common form, such as in column headers of flat files. They can be insufficient for a proper understanding of what is measured and often gather information of different natures. In order to solve these ambiguities, a semantic decomposition work is carried out for each variable, driven by the model and the terms of the ontology used. In our case, the principle is based on the identification of a characteristic measured according to a standard (a unit) for the observation of an entity, itself potentially contextualised by one or several observations of other entities.</p><p>As an illustration, the variable named "Ammonia nitrogen" in a physical chemistry database on lake data is decomposed as shown in Figure <ref type="figure" target="#fig_2">3</ref>. In addition to this semantic decomposition that will be used for annotation, the usual name of each variable is supplied to facilitate communication and to easily respond to certain uses such providing column headings of data files, metadata fields or drop-down lists of available variables. To do this, two naming standards have been added to the ontology as classes, one for the usual variable names (Anaee-France variableNamingStandard), the other for variable categories (Anaee-France variable category naming standard). Each of these classes contains a list of items as individuals. In the previous example of the variable "Ammonia nitrogen", the standardised usual name retained is "dissolved ammonium nitrogen mass concentration" and it is associated with the standard category "physical chemistry". A dedicated property named "hasVariableContext" has also been added to the ontology and is used to associate the result of the previous semantic decomposition with the standardised usual name of the variable and with one or several standard categories (Figure <ref type="figure" target="#fig_2">3</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3.">Generic Graph Models</head><p>The complete modelling of a graph for a measured variable implies adding other contextual elements such as spatial, temporal or location information through the reflexive property "hasContext" of the ontology (Figure <ref type="figure" target="#fig_3">4</ref>). This results in a rich graph, more or less complex depending on the variables and the use cases. The modelling of a graph for one variable is most often suitable to other variables. Indeed if the values of the modelled information are different from one variable to another, their natures and their structures are common to several variables. Sets of variables can thus share the same graph structure defining a common pattern called "graph model". Thus, the graph initially constructed for one variable can be generalised to multiple ones by introducing a set of elements whose values vary according to the variable being processed. These so-called "dynamic elements" are the usual name of the variable, the entities observed, the characteristics, the measurement standards, the entities of the near context observations (e.g. matrix) or the thematic categories of the variables. The resulting graph model is then instantiated as a specific graph for each of the related variables. Whatever the variables, the graphs systematically contain similar graph structures from the semantic decomposition of the variables and their standardised naming. The common part shared by all the graph models is presented in Figure <ref type="figure" target="#fig_4">5</ref>. The dynamic elements used for the instantiation of a graph for a given variable can be single (standardised variable name) or multiple (entities og the context observations and categories). Their values, resulting from the semantic analysis of the variables, are provided in a dedicated file, csv format, where each line corresponds to a single variable (Table <ref type="table" target="#tab_0">1</ref>). The instantiation of the graphs, for each variable, is then delegated to the first pipeline, responsible for the semantic annotation.</p><p>Thus, this generalisation of graphs through patterns called "graph models" and the use of an input parameters csv file makes it possible (i) to make the modelling effort more generic and (ii) to automate the instantiation of graphs for each variable to be processed.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Implementation for Planktonic Biodiversity Data from Lakes</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.">Variable Description</head><p>An application of the workflow described in section 2 was conducted on the Observatory on LAkes (OLA) database <ref type="bibr" target="#b1">[2]</ref> by mobilising the total phytoplankton ("biovolume") and the zooplankton per taxon ("sedimented volume") as biodiversity data, and "water temperature" and "dissolved orthophosphorus" as environmental data. These variables were processed for the Lake Geneva data and the period 1974-2004. The semantic decomposition of the different variables involved was carried out with the domain experts and produced the file of Table <ref type="table">2</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 2</head><p>Semantic decomposition and standard naming of chosen variables for planktonic biodiversity data. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Standard</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2.">Variable Modelling</head><p>Following the modelling principles described in section 3, two models generated for this implementation are illustrated below. The first (Figure <ref type="figure" target="#fig_5">6</ref>) concerns the physico-chemical environment variable orthophosphorus and is an instantiated graph for this variable, automatically generated by a pipeline from a graph model. The second (Figure <ref type="figure" target="#fig_6">7</ref>) is a graph model applicable to biodiversity data and used in this implementation for phyto and zooplankton data.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.3.">Data Sets and Metadata Records</head><p>Defined on the perimeter 'Lake Geneva x [water temperature, orthophosphorus, zooplankton, phytoplankton] x ', a dataset and a metadata record were generated and published by the AnaEE workflow (doi:10.15454/XZWVM8). The discovery and exploitation metadata respectively present in the GeoDCAT file (Box 1) and the NetCDF file header (Box 2) were automatically filled in with data from the database and semantic information from the graph model and from the ontology.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Box 1</head><p>Extract from the GeoDCAT metadata record.</p><p>&lt;dcterms:description xml:lang="en"&gt;The data set was produced from experimentation(s) from the network(s) SOERE OLA on the site(s) Lake Geneva in the ecosystem(s) lake.  Figure <ref type="figure" target="#fig_7">8</ref> illustrates the seasonality of phosphorus and also the drop in concentrations induced by restoration measures aimed at limiting P inputs to the lake. The decrease in zooplankton means less grazing pressure on phytoplankton and therefore less effective regulation <ref type="bibr" target="#b5">[6]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Conclusion and Perspectives</head><p>Although of main importance for interoperability, the implementation of the semantic characterisation of data and metadata is a difficult task to implement as it is very costly in terms of shared technical and semantic resources and in terms of annotation of the data and their acquisition contexts.</p><p>The work presented and carried out in the framework of AnaEE and the ENVRI-plus/FAIR projects aims to facilitate this implementation. The genericity of OBOE allows a modelling of the whole system. Its Entity-Observation-Characteristics model also allows an alignment with other ontologies based on entity property relationships (O&amp;Ḿ, SOSA...). The convergence of these models is the subject of ongoing work in the framework of the RDA I-Adopt WG. The challenge here is not only to develop interoperability within the theoretical communities but also between the domains.</p><p>The enrichment of OBOE modelling, for components concerning people and sensors via specific ontologies (FOAF, SOS, etc.) is a short-term prospect. The PROV ontology (PROV-O) will be used to model provenance elements from data acquisition to dataset publication, triples being generated by the developed pipelines, from a dedicated graph.</p><p>Although the deployment of the AnaEE workflow is currently limited to French experimental platforms and does not cover all variables, the strategy being pursued is to implement it as systematically as possible, based on the experience gained.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Semantic workflow for the management and valorisation of the data produced by the AnaEE platforms</figDesc><graphic coords="3,86.20,72.00,302.40,175.65" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Figure 2 :</head><label>2</label><figDesc>Figure 2: The core classes and properties of the Extensible Observation Ontology (OBOE), from Madin et al. 2007[5]).</figDesc><graphic coords="3,86.20,439.23,301.05,99.20" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Figure 3 :</head><label>3</label><figDesc>Figure 3: semantic decomposition of "Ammonia nitrogen" variable according to the ontology model and terms (upper left) with corresponding representation as a graph (upper right). The graph also provides the correspondence to the usual variable and category(ies) names (bottom box).</figDesc><graphic coords="4,86.20,229.58,436.45,255.40" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_3"><head>Figure 4 :</head><label>4</label><figDesc>Figure 4 : Complete graph modelling overview for the "Ammonia nitrogen" example.</figDesc><graphic coords="5,86.20,109.95,436.75,198.25" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_4"><head>Figure 5 :</head><label>5</label><figDesc>Figure 5: Graph model resulting from the generalisation, applicable to all the variables.</figDesc><graphic coords="6,86.20,72.00,437.30,381.70" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_5"><head>Figure 6 :</head><label>6</label><figDesc>Figure 6: Instance of the physico-chemical data graph model for the variable dissolved orthophosphorus mass concentration.</figDesc><graphic coords="8,86.20,72.00,439.35,611.45" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_6"><head>Figure 7 :</head><label>7</label><figDesc>Figure 7: Graph model for plankton variables. The part specific to the experimental context is not detailed here.</figDesc><graphic coords="9,86.20,72.00,434.10,363.55" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_7"><head>Figure 8 :</head><label>8</label><figDesc>Figure 8: Relationship between orthophosphorus and phytoplankton biovolume (all species) in Lake Geneva, based on semantic data produced by the workflow described in present document.</figDesc><graphic coords="10,86.20,479.57,436.50,200.95" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_0"><head>Table 1</head><label>1</label><figDesc>Semantic decomposition and standard naming of variables. In blue, the unique elements and in orange, the potentially multiple elements.</figDesc><table><row><cell>Standard Variable Name</cell><cell>Category (ies)</cell><cell>Context(s)</cell><cell>Entity</cell><cell>Characteristic</cell><cell>Standard Measurement</cell></row><row><cell>Dissolved Ammonium Nitrogen Mass Concentration</cell><cell>Physical Chemistry</cell><cell>Water, Solutes, Ammonium</cell><cell>Nitrogen</cell><cell>Mass Concentration</cell><cell>Milligram Per Liter</cell></row><row><cell>Calcium Mass Concentration</cell><cell>Physical Chemistry</cell><cell>Water</cell><cell>Calcium</cell><cell>Mass Concentration</cell><cell>Milligram Per Liter</cell></row><row><cell>WaterPH</cell><cell>Physical Chemistry</cell><cell></cell><cell>Water</cell><cell>pH</cell><cell>pH unit</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>MicroSquare</cell></row><row><cell>Biovolume</cell><cell>Biodiversity</cell><cell>Water</cell><cell>Zooplankton</cell><cell>Biovolume</cell><cell>Meter Per</cell></row><row><cell></cell><cell></cell><cell></cell><cell></cell><cell></cell><cell>Millilitre</cell></row><row><cell>...</cell><cell>...</cell><cell>...</cell><cell>...</cell><cell>...</cell><cell>...</cell></row></table></figure>
		</body>
		<back>

			<div type="acknowledgement">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.">Acknowledgements</head><p>The work was co-funded by AnaEE-France, French program "Investissements d'Avenir" (ANR-11-INBS-0001), the European Horizon 2020 ENVRIplus project (No 654182) and ENVRI-FAIR project (No 824068) and the D2KAB project (ANR-18-CE23-0017). It notably relies the contribution of the AnaEE semantic working group (*) who develops the vocabularies: i) the anaeeThes thesaurus (http://agroportal.lirmm.fr/ontologies/ANAEETHES) ii) the anaee ontology extension of OBOE (ongoing publication). We are also grateful to the INRAE-CARRTEL technical and scientific team that collected and provided the data from the Geneva Lake Observatory. (*) Pichot C. (1), Callou C. (2), Chanzy A. (3), Clastre P. (1), Clavreul A. (3), El-Hamadry M. (4), Evtimova M. (1), Jaillet B. (1), Lafolie F. (3), Le Gaillard J.-F. (5), Martin C. (2), Massol F. (5), Maurice D. (6), Moitrier N. (3), Monet G. (7), Raynal H. (4), Schellenberger A. (8), Yahiaoui R. (8), Aïvayan E. (3), Beudez N. (3), Léturgie A. (1) 1. INRAE URFM, 2. CNRS UMS BBEES, 3. INRAE UMR EMMAH, 4. INRAE MIAT, 5. UMS CEREEP-Ecotron, 6. INRAE UMR SILVA, 7. INRAE UMR CARRTEL, 8. INRAE INFOSOL.</p></div>
			</div>

			<div type="annex">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Box 2</head><p>Extract from the NetCDF file header for water temperature and orthophosphorus concentration. Water temperature ('Var0') is expressed in degree Celcius and provides data collected in one experimental plot ('Shl2Platform', 46.45°N, 6.59°E ) at 25 depths (Dim0) expressed in meter and for 559 dates (Dim1) Var0:description = " temperature of water in degree Celsius" ; Var0:name_of_experimental_site_in_text = "Léman" ; Var0:name_of_variable_in_text = "Temperature" ; Var0:name_of_ecosystem_type_in_Anaee-France_ecosystem_type_naming_standard = "http://opendata.inra.fr/anaeeOnto#Lake" ; Var0:name_of_experimental_network_in_Anaee-France_experimental_network_naming_standard = "http://opendata.inra.fr/anaeeOnto#SoereOla" ; Var0:name_of_experimental_plot_in_Anaee-France_experimental_plot_naming_standard = "http://opendata.inra.fr/anaeeOnto#Shl2Platform" ; Var0:name_of_experimental_site_in_Anaee-France_experimental_site_naming_standard = "http://opendata.inra.fr/anaeeOnto#LakeGeneva" ; Var0:name_of_variable_in_Anaee-France_variable_naming_standard = "http://opendata.inra.fr/anaeeOnto#WaterTemperature" ; Var0:latitude_of_Waypoint_in_decimal_degree = "46.453457" ; Var0:longitude_of_Waypoint_in_decimal_degree = "6.5942335" ; double Var1(Var1Dim0, Var1Dim1) ;</p><p>../.. // global attributes:</p><p>:lineage = "The dataset was generated, formatted and published using AnaEE semantic services and vocabularies" ; data: Var0Dim0 = "0.0", "10.0", "100.0", "15.0", "150.0", "2.5", "20.0", "200.0", "225.0", "25.0", "250.0", "275.0", "280.0", "285.0", "290.0", "295.0", "30.0", "300.0", "305.0", "309.0", "35.0", "40.0", "5.0", "50.0", "7.5" ; Var0Dim1 = "1974-01-14", "1974-02-18", "1974-03-18", "1974-04-22", "1974-05-13", "1974-06-17", "1974-07-15", "1974-08-19", "1974-09-16",</p></div>			</div>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<analytic>
		<title level="a" type="main">How to Integrate Experimental Research Approaches in Ecological and Environmental Studies: AnaEE France as an Example</title>
		<author>
			<persName><forename type="first">J</forename><surname>Clobert</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Chanzy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">F</forename><surname>Le Galliard</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Chabbi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Greiveldinger</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Caquet</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Loreau</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Mougin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Pichot</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Roy</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Saint-André</surname></persName>
		</author>
		<idno type="DOI">10.3389/fevo.2018.00043,hal-01773144</idno>
	</analytic>
	<monogr>
		<title level="j">Front. Ecol. Evol</title>
		<imprint>
			<biblScope unit="volume">6</biblScope>
			<biblScope unit="page">43</biblScope>
			<date type="published" when="2018">2018</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">The Observatory on LAkes (OLA) database: Sixty years of environmental data accessible to the public</title>
		<author>
			<persName><forename type="first">F</forename><surname>Rimet</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Anneville</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Barbet</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Chardon</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Crepin</surname></persName>
		</author>
		<idno type="DOI">10.4081/jlimnol.2020.1944,hal-02916312</idno>
	</analytic>
	<monogr>
		<title level="j">Journal of Limnology</title>
		<imprint>
			<biblScope unit="volume">79</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page" from="164" to="178" />
			<date type="published" when="2020">2020</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Pipelines for semantic annotation and FAIR data production</title>
		<author>
			<persName><forename type="first">C</forename><surname>Pichot</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Maurice</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Clastre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Jaillet</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><surname>Yahiaoui</surname></persName>
		</author>
		<idno>hal- 03234314</idno>
	</analytic>
	<monogr>
		<title level="m">RDA 17th Plenary Meeting</title>
				<meeting><address><addrLine>Edinburgh, United Kingdom</addrLine></address></meeting>
		<imprint>
			<date type="published" when="2021-04">Apr 2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<monogr>
		<title level="m" type="main">OBOE: the Extensible Observation Ontology</title>
		<author>
			<persName><forename type="first">M</forename><surname>Schildhauer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">B</forename><surname>Jones</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Bowers</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Madin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Krivov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Pennington</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Villa</surname></persName>
		</author>
		<author>
			<persName><forename type="first">B</forename><surname>Leinfelder</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Jones</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>O'brien</surname></persName>
		</author>
		<idno type="DOI">10.5063/F1125R0F</idno>
		<imprint>
			<date type="published" when="2016">2016</date>
			<publisher>KNB Data Repository</publisher>
		</imprint>
	</monogr>
	<note>version 1.2</note>
</biblStruct>

<biblStruct xml:id="b4">
	<analytic>
		<title level="a" type="main">An ontology for describing and synthesizing ecological observation data</title>
		<author>
			<persName><forename type="first">J</forename><surname>Madin</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Bowers</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Schildhauer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Krivov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Pennington</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Villa</surname></persName>
		</author>
		<idno type="DOI">10.1016/j.ecoinf.2007.05.004</idno>
	</analytic>
	<monogr>
		<title level="j">Ecological Informatics</title>
		<imprint>
			<biblScope unit="volume">2</biblScope>
			<biblScope unit="page" from="279" to="296" />
			<date type="published" when="2007">2007</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">The paradox of reoligotrophication: the role of bottom-up versus top-down controls on the phytoplankton community</title>
		<author>
			<persName><forename type="first">O</forename><surname>Anneville</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Chang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Dur</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Souissi</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Rimet</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Hsieh</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Oikos</title>
		<imprint>
			<biblScope unit="volume">128</biblScope>
			<biblScope unit="issue">11</biblScope>
			<biblScope unit="page" from="1666" to="1677" />
			<date type="published" when="2019">2019</date>
		</imprint>
	</monogr>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
