Fully automated non-targeted GC-MS data analysis

Abstract

Non-targeted analysis is applied in many different domains of analytical chemistry such as metabolomics, environmental and food analysis. In contrast to targeted analysis, non-targeted approaches take information of known and unknown compounds into account, are inherently more comprehensive and give a more holistic representation of the sample composition. 

Besides chromatographic techniques coupled to high resolution mass spectrometry such as LC-HRMS, gas chromatography with unit resolution mass spectrometry is still regularly utilized for non-targeted profiling or fingerprinting. This is mainly due to high separation power of GC and a wide availability and low costs of quadrupole mass spectrometers. 

Although several non-targeted approaches have been developed, data processing still remains a serious bottleneck. Baseline correction, feature detection, and retention time alignment can be prone to errors and time-consuming manual corrections are often necessary. We therefore developed an automated strategy to non-targeted GC-MS data avoiding feature detection and retention time alignment. The novel automated approach includes segmentation of chromatograms along the retention time axis, multiway decomposition of transformed segments followed by a supervised machine learning pipeline based on gradient boosted tree classification on the decomposed tensor [1, 2]. 

In order to make this novel data analysis strategy available to scientists without programming background, we developed a convenient browser based application. For the here presented interactive browser application the open source Python packages Bokeh and HoloViews were used. The application will be online freely available soon. 

[1] J. Vestner, G. de Revel, S. Krieger-Weber, D. Rauhut, M. du Toit, A. de Villiers, Toward automated chromatographic fingerprinting: A non-alignment approach to gas chromatography mass spectrometry data. Acta Chimica Acta 911 (2016) 42-58 
[2] K. Sirén, U. Fischer, J. Vestner, Automated supervised learning pipeline for non-targeted GC-MS data analysis. Analytica Chimica Acta: X 1 (2019) 100005

DOI:

Publication date: June 19, 2020

Issue: OENO IVAS 2019

Type: Article

Authors

Jochen Vestner, Kimmo Sirén, Pierre Le Brun, Ulrich Fischer

Institute for Viticulture and Oenology, DLR Rheinpfalz, Breitenweg 71, D-67435 Neustadt, Germany
Institut National Supérieur des Sciences Agronomiques de l’Alimentation et de l’ Environnement, Agrosup Dijon, 6 boulevard Docteur Petitjean, 21000 Dijon, France
Department of Chemistry, University of Kaiserslautern, Erwin-Schroedinger-Strasse 52, D-67663 Kaiserslautern

Contact the author

Keywords

metabolomics, non-targeted, GC-MS, exploratory data analysis 

Tags

IVES Conference Series | OENO IVAS 2019

Citation

Related articles…

An analytical framework to site-specifically study climate influence on grapevine involving the functional and Bayesian exploration of farm data time series synchronized using an eGDD thermal index

Climate influence on grapevine physiology is prevalent and this influence is only expected to increase with climate change. Although governed by a general determinism, climate influence on grapevine physiology may present variations according to the terroir. In addition, these site-specific differences are likely to be enhanced when climate influence is studied using farm data. Indeed, farm data integrate additional sources of variation such as a varying representativity of the conditions actually experienced in the field. Nevertheless, there is a real challenge in valuing farm data to enable grape growers to understand their own terroir and consequently adapt their practices to the local conditions. In such a context, this article proposes a framework to site-specifically study climate influence on grapevine physiology using farm data. It focuses on improving the analysis of time series of weather data. The analytical framework includes the synchronization of time series using site-specific thermal indices computed with an original method called Extended Growing Degree Days (eGDD). Synchronized time series are then analyzed using a Bayesian functional Linear regression with Sparse Steps functions (BLiSS) in order to detect site-specific periods of strong climate influence on yield development. The article focuses on temperature and rain influence on grape yield development as a case study. It uses data from three commercial vineyards respectively situated in the Bordeaux region (France), California (USA) and Israel. For all vineyards, common periods of climate influence on yield development were found. They corresponded to already known periods, for example around veraison of the year before harvest. However, the periods differed in their precise timing (e.g. before, around or after veraison), duration and correlation direction with yield. Other periods were found for only one or two vineyards and/or were not referred to in literature, for example during the winter before harvest.

Postveraison shoot trimming in Tannat and Merlot: preliminary results on yield components, plant balance and berry composition

There is currently a trend towards the production of wines with low alcohol content. To achieve this, grapes with low sugar content must be used. There are techniques at the vineyard level that can delay ripening and avoid excessive sugar accumulation without, a priori, affecting the final polyphenol content. Postveraison shoot trimming (PVST) is experimentally evaluated for these purposes, but its impact under Uruguayan climatic conditions with high interannual variability is not known. The aim of this work is to assess the PVST in Tannat and Merlot cultivars and their impact on yield components, plant balance and berry primary composition. In this study, two commercial vineyards of 10 years old Tannat and Merlot (grafted on SO4) at Canelones Department were selected. During the 2020-201 growing season, grapevines were submitted to PVST when grapes reached 15º Brix. In a randomized block, trimmed (T) and control (C) plants were evaluated with three repetitions each cultivar. Evaluation of the evolution of primary berry composition during ripening, measurement of yield components and plant balance were performed. For both cultivars, PVST did not affect yield components. Merlot reached 5.4 kg per plant and Tannat 7.1 kg, with not statistical significance between treatments. However, statistical differences were observed in terms of plant balance. In Merlot Ravaz Index reached a difference of 5.3 (12.0 in T and 6.7 in C) meanwhile Tannat reached 3.5 of statistical difference (13.7 in T and 10.2 in C). The tendency to imbalance for the treated plants had an impact on the final grape composition. Merlot grapes showed statistical difference in final total acidity (0.3 g of difference between treatments) while treatments impact final sugar content on Tannat grapes (10.0 g of difference between treatments). Further studies are needed to assess the impact of different canopy management techniques in our conditions.

Protected Designation of Origin (D.P.O.) Valdepeñas: classification and map of soils

The objective of the work described here is the elaboration of a map of the different types of vineyard soils that to guide the famers in the choice of the most productive vine rootstocks and varieties. 90 vineyard soils profiles were analysed in the entire territory of the Origen Denominations of Valdepeñas. The sampling was carried out in 2018 (June to October) by making a sampling grid, followed by photointerpretation and control in the field. The studied soils can be grouped into 9 different soil types (according to FAO 2006 classification): Leptosols, Regosols, Fluvisols, Gleysols, Cambisols, Calcisols, Luvisols and Anthrosols. A map showing the soil distribution with different type of soils has been made with the ArcGIS program. Regarding to the choice of rootstock, Calcisoles are soils with a high active limestone content, so the rootstocks used in these soils must be resistant to this parameter; Luvisols are deep soils with high clay content, so they will support vigorous rootstocks. Because the cartographic units are composed of two or more subgroups, with are associated in variable proportions, 9 different soil associations have been established; Unit 1: Leptosols, Cambisols and Luvisols (80%, 15% and 5% respectively); Unit 2: Cambisols with Regosols and Luvisols (40%, 30% and 30% respectively); Unit 3: Cambisols and Gleysols with Regosols (40%, 40% and 20% respectively); Unit 4: Regosols with Cambisols, Leptosols and Calcisols (40%, 30%, 15% and 15% respectively); Unit 5: Cambisols, Leptosols, Calcisols and Regosols (25% each of them); Unit 6: Luvisols with Cambisol and Calcisols (80%, 10% and 10% respectively); Unit 7: Luvisols and Calcisols with Cambisols (40%, 40% and 20% respectively); Unit 8: Calcisols with, Cambisols and Luvisols (80%, 10% and 10% respectively); Unit 9: Anthrosols. These study allow to elaborate the first map of vineyard soils of this Protected Designation of Origin in Castilla-La Mancha.

Legacy of land-cover changes on soil erosion and microbiology in Burgundian vineyards

Soils in vineyards are recognized as complex agrosystems whose characteristics reflect complex interactions between natural factors (lithology, climate, slope, biodiversity) and human activities. To date, most of the unknown lies in an incomplete understanding of soil ecosystems, and specifically in the microbial biodiversity even though soil microbiota is involved in many key functions, such as nutrient cycling and carbon sequestration. Soil biological properties are indicative of soil quality. Therefore, understanding how soil communities are related to soil ecosystem functioning is becoming an essential issue for soil strategy conservation. Here, we propose to assess the importance of land-cover history on the present-day microbiological and physico-chemical properties. The studied area was selected in the Burgundian vineyards (Pernand-Vergelesses, Burgundy, France) where land occupation has been reconstructed over the last 40 years. Soil samples were collected in five areas reflecting various land cover history (forest, vineyards, shifting from forest to vineyards). For each area, physico-chemical parameters (pH, C, N, P, grain size) were measured and DNA was extracted to characterize the abundance and diversity of microbial communities. The obtained results show significant differences in the five areas suggesting that present-day microbial molecular biomass and bacterial taxonomic is partly inherited from past land occupation. Over longer period of time, such study of land-uses legacies may help to better assess ecosystem recovery and the impact of management practices for a better soil quality and vineyards sustainability.

Towards a regional mapping of vine water status based on crowdsourcing observations

Monitoring vine water status is a major challenge for vineyard management because it influences both yield and harvest quality. It is also a challenge at the territorial scale for identifying periods of high water restriction or zones regularly impacted by water stress. This information is of major importance for defining collective strategies, anticipating harvest logistic or applying for irrigation authorisation. At this spatial scale, existing tools and methods for monitoring vine water status are few and often require strong assumptions (e.g. water balance model). This paper proposes to consider a collaborative collection of observations by winegrowers and wine industry stakeholders (crowdsourcing) as an interesting alternative. Indeed, it allows the collection of a large number of field observations while pooling the collection effort. However, the feasibility of such a project and its interest in monitoring vine water status at regional scale has never been tested.

The objective of this article is to explore the possibility of making a regional map of vine water status based on crowdsourcing observations. It is based on the study of the free mobile application ApeX-Vigne, which allows the collection of observations about vine shoot growth. This information is easy to collect and can be considered, under certain conditions, as a proxy for vine water status. This article presents the first results obtained from the nearly 18,000 observations collected by winegrowers and wine industry stakeholders during 2019, 2020 and 2021 seasons. It presents the vine shoot growth maps obtained at regional scale and their evolution over the three vintages studied. It also proposes an analysis of the factors that favoured the number of observations collected and those that favoured their quality. These results open up new perspectives for monitoring vine water status at a regional scale but above they provide references for other crowdsourcing projects in viticulture.