Trends in crossplots of compositional data from different analyses are statistically valid
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
6 sources for · 0 against
Peer-reviewed literature demonstrates that compositional data crossplots and ternary diagrams can yield valid statistical trends when appropriate methods, such as log-ratio transformations or specialized closed-data considerations, are applied to overcome spurious correlation.
Abstract Geochemical data are typically reported as compositions, in the form of some proportions such as weight percents, parts per million, etc., subject to a constant sum (e.g. 100%, 1,000,000 ppm). This latter implies that such data are “closed”; that is, for a composition of D -components, only D − 1 components are required. The statistical analysis of compositional data has been a major issue for more than 100 years. The problem of spurious correlation, introduced by Karl Pearson in 1897, affects all data measuring parts of some whole, which are by definition, constrained; and such type of measurements are present in all fields of geochemical research. The use of the log-ratio transform was introduced by John Aitchison to overcome these constraints by opening the data into the real number space, within which standard statistical methods can be applied. However, many statisticians and users of statistics in the field of geochemistry are unaware of the problems affecting compositional data, as well as solutions that overcome these problems. A look into the ISI Web of Science and Scopus databases shows that most papers where compositional data are the core of a geochemical research continue to ignore methods to correctly manage constrained data. A key question is how we can demonstrate that the interpretation of the behaviour of chemical species in natural environment and in geochemical processes is improved when the compositional constraint of geochemical data is taken into account through the use of new methods. In order to achieve this aim, this special issue of the Journal of Geochemical Exploration focuses on the correct statistical analysis of compositional data. Applications in exploration, monitoring and environments by considering several geological matrices are presented and discussed illustrating that several paths can be followed to understand how geochemical processes work.
Abstract Ore deposits usually consist of ore materials with different discrete (e.g. rock and alteration types) and continuous (e.g. geochemical and mineral composition) features. Financial feasibility studies are highly dependent on the modelling of these features and their associated joint uncertainties. Few geostatistical techniques have been developed for the joint modelling of high-dimensional mixed data (continuous and categorical) or constrained data, such as compositional data. The compositional nature of the mineral and geochemical data induces several challenges for multivariate geostatistical techniques, because such data carry relative information and are known for spurious statistical and spatial correlation effects. This paper investigates the application of the direct sampling algorithm for joint modelling of compositional and categorical data. In some mining projects the amount of available data may be enormous in some parts of the deposit and if the density of measurements is sufficient, multivariate geospatial patterns can be derived from that data and be simulated (without model inference) at other undersampled areas of the deposit with similar characteristics. In this context, the direct sampling multiple-point simulation method can be implemented for this reconstruction process. The compositional nature of the data is addressed via implementing an isometric log-ratio transformation. The approach is illustrated through two case studies, one synthetic and one real. The accuracy of the results is checked against a set of validation data, revealing the potential of the proposed methodology for joint modelling of compositional and categorical information. The direct sampling technique can be considered as a smart move to assess the future risk and uncertainty of a resource by making use of all the information hidden within the early data.
Shale samples from source rocks of the Upper Triassic Yanchang Formation (Chang 9 member) in the Ansai area, Ordos Basin, North China, were analyzed using gas chromatography - mass spectrometry (GC-MS) to investigate the distribution, abundance, and enrichment mechanisms of rearranged hopanes. Four rearranged hopane series were detected, with all four present simultaneously in individual samples. Analysis of the C₃₀ hopane series (regular C₃₀H, diahopane C₃₀D, and neohopane C₃₀E) using a ternary diagram revealed a distinct linear trend, demonstrating a systematic, inverse relationship between the abundance of regular hopane and the combined abundance of its rearranged counterparts. These results provide strong evidence that C₃₀D and C₃₀E in the Chang 9 shales are diagenetic products derived from C₃₀H, sharing a common biological precursor. Both diasteranes and regular steranes with the ββ configuration were correlated positively in abundance with rearranged hopanes, further supporting a common origin linked to specific organism assemblages rather than widespread organisms. Samples deposited under highly saline, suboxic sedimentary environments displayed relatively high abundances of rearranged hopanes, indicating the critical role of depositional conditions in their enrichment. Multi-proxy analysis revealed a complex, non-linear control of thermal maturity on rearranged hopane abundance. The C₃₀ Rearranged Hopane Index showed statistically significant positive correlations with multiple maturity parameters (including sterane and hopane isomerization ratios), indicating maturity as a primary driver in the early oil window. However, this trend diverged at higher maturity levels, suggesting that other factors, such as the catalytic activity of the mineral matrix, become dominant. Our findings establish a robust biomarker-based framework for interpreting oil-source correlations and informing petroleum exploration in the Ordos Basin, particularly for the Chang 9 member s
The C₃₀ Rearranged Hopane Index showed statistically significant positive correlations with multiple maturity parameters (including sterane and hopane isomerization ratios), indicating maturity as a primary driver in the early oil window. However, this trend diverged at higher maturity levels, suggesting that other factors, such as the catalytic activity of the mineral matrix, become dominant. Our findings establish a robust biomarker-based framework for interpreting oil-source correlations and informing petroleum exploration in the Ordos Basin, particularly for the Chang 9 member source rocks.
3.4 Statistical analysis All statistical analyses were performed using Microsoft Excel 365 (Microsoft Corporation, Redmond, WA, USA) with the Real Statistics Resource Pack add-in (version 9.5.5). To avoid the statistical closure problem associated with compositional data, a ternary diagram was used to visualize the relative proportions of C₃₀ hopanes and a C₃₀ Rearranged Hopane Index (RHI) for quantitative analysis. Spearman’s rank correlation analysis (r s ), a non-parametric method that is robust to small sample sizes and non-normally distributed data, was employed to assess monotonic relationships between key biomarker parameters.
Furthermore, the data points spread across the oxidizing-to-reducing trend lines, indicating that this organic matter was deposited under fluctuating, oxygen-limited conditions. Interpretive bands modified after Shanmugam (1985) [ 38 ] and Peters et al. (2005) [ 33 ]. Sample PE312 is explicitly labeled for clarity, and sample X762 (red symbol) is highlighted. However, a more precise interpretation is provided by the Pr/ n -C 17 versus Ph/ n -C 18 cross-plot ( Fig 3b ), a diagnostic tool widely used for its ability to effectively deconvolve the interconnected effects of organic matter source and redox conditions [ 38 ].
In Fig 3b , this ambiguity is clearly resolved, as PE312 is located at the end of this trend, plotting clearly within the ‘Terrestrial organic matter’ field and in the ‘Oxidizing’ position. This
The conventional view holds that rearranged hopanes share identical biological precursors with regular hopanes, originating primarily from bacteriohopanetetrol [ 10 , 42 ]. This implies a precursor–product relationship in which regular hopanes are converted into rearranged hopanes during diagenesis [ 10 , 43 ]. To investigate this relationship in the Chang 9 samples while avoiding the statistical pitfalls of compositional data closure, we visualized the relative proportions of C₃₀ diahopane (C₃₀D), C₃₀ neohopane (C₃₀E), and regular C₃₀ hopane (C₃₀H) using a ternary diagram ( Fig 9a ).
The anomalous sample X762 (red symbol) is highlighted to distinguish it from the main group (blue symbols). The data points form a clear linear trend extending from the C₃₀H apex towards the C₃₀D–C₃₀E baseline. This distribution strongly supports a systematic conversion process in which the abundance of C₃₀H decreases as those of C₃₀D and C₃₀E increase. To quantify the extent of this conversion, we defined a C₃₀ Rearranged Hopane Index (RHI) as (C₃₀D + C₃₀E)/ (C₃₀D + C₃₀E + C₃₀H). As shown in Fig 9b , there was a mathematically defined inverse relationship between the RHI and the relative proportion of regular C₃₀ hopane (%C₃₀H).
To robustly evaluate the impact of maturity, and in response to the valid concern that using single maturity proxies or burial depth across different wells can be misleading, we conducted a multi-proxy analysis. We assessed the relationship between the C₃₀ RHI and four different maturity parameters derived from aromatic, triterpane, and sterane compounds ( Fig 13 ). The outlier sample X762, which exhibited the highest RHI, was excluded from the correlation analysis to assess the general trend of the majority of the samples. 10.1371/journal.pone.0337076.g013 Fig 13 Relationships between rearranged hopane abundance (RHI) and multiple thermal maturity proxies for the Chang 9 shales.
Potentially toxic elements (PTEs) such as arsenic (As) and mercury (Hg) are among the most critical pollutants globally, threatening ecosystem integrity and human health. The Trimpancho mining system in the Iberian Pyrite Belt (W Spain) is one such hotspot, where centuries of activity have left a legacy of acid mine drainage and heavy metal dispersion. This study employs an integrated compositional, probabilistic, and spatial modeling framework to characterize and map contamination dynamics in this area with quantified uncertainty. A total of 31 water samples were collected during 2022 and 2023 from surface streams and tributaries. Concentration data were transformed using isometric log-ratio (ilr) techniques to preserve their compositional nature and avoid spurious correlations. Bayesian Networks (BNs), combined with information-theoretic metrics, were then applied to identify latent geochemical contamination patterns and quantify both aleatory and epistemic uncertainties. The key drivers identified were incorporated into a co-kriging framework, enabling spatial interpolation that accounted for over 90% of total variance and reduced epistemic uncertainty by 22.7% compared to raw-data models. The resulting spatial–temporal maps revealed distinct As–Hg contamination signatures, influenced by hydrological variability and mining legacy sources. In conclusion, this integrated approach provides a robust, uncertainty-aware methodology for detecting, interpreting, and mapping contamination patterns, offering actionable insights for environmental risk assessment and remediation planning in mining-impacted watersheds.
Concepts of null correlation for r-compositions are discussed in this chapter, following the methodology developed by J. Aitchison for the statistical analysis of compositional data. This will be combined with G. Matheron’s theory of regionalized variables. These concepts are to be understood in the sense of absence de correlation différée (absence of deferred correlation), as defined by Matheron (1965). Concepts of null correlation are important not only for spatial-structure analysis of r-compositions, but also for simulation of phenomena that can be described by the use of r-compositions. The intrinsic analogue to the definitions of null correlation in the secondorder stationary case is carried out in parallel in this chapter because the relation between them is of special interest. All of the following concepts depend in general on the length of the vector h and also on its direction and sign, that is, they can be defined depending on the length, or set of lengths, and the direction, or set of directions, of h, or both. Therefore, as in Chapter 3, statements will be made for h Î H, where H stands for a set of vectors with specified range of directions and range of lengths. Here, for example, H may contain all possible directions and lengths for h, except h = 0; in this case, statements will be valid only in a spatial sense, but not in a standard nonspatial sense.
In Chapter 6 we introduce additional aspects of the geostatistical approach presented in the preceding chapters that were not necessary for its theoretical development, but that are essential for the practical application of the method to compositional data. We discuss how to treat zeros in compositional data sets; how to model the required cross-covariances; how to compute expected values and estimation variances for the original, constrained variables; and how to build and interpret confidence intervals for estimated values. As mentioned in Section 2.1, data sets with many zeros are as troublesome in compositional analysis as they are in standard multivariate analysis. In our approach, the additional restriction for compositional data is that zero values are not admissible for modeling. The justification for this restriction can be given using arithmetic arguments. A transformation that uses logarithms cannot be performed on zero values. This is the case for the logratio transformation that leads to the definition of an additive logistic normal distribution, as introduced by Aitchison (1986, p. 113). It is also the case for the additive logistic skew-normal distribution defined in Mateu-Figueras et al. (1998), following previous results by Azzalini and Dalla Valle (1996). The centered logratio transformation and the family of multivariate Box-Cox transformations discussed in Andrews et al. (1971), Rayens and Srinivasan (1991), and Barceló- Vidal (1996) also call for the restriction of zero values. This restriction is certainly a wellspring of discussion, albeit surprisingly so, as nobody would complain about eliminating zeros either by simple suppression of samples or by substitution with reasonable values when dealing with a sample from a lognormal distribution in the univariate case. Recall that the logarithm of zero is undefined and the sample space of the lognormal distribution is the positive real line, excluding the origin. In order to present our position on how to deal with zeros as clearly as possible, let us assume that only one of our components has zeros in some of the samples. Those cases where more than one variable is affected can be analyzed by methods described below.
Everything we examined (6)
This check searched the claim as stated. It did not run a separate search for evidence against it.