Package {Compositionalcln}


Type: Package
Title: Modelling Compositional Data with Zero Values
Version: 1.1
Date: 2026-10-06
Author: Michail Tsagris [aut, cre]
Maintainer: Michail Tsagris <mtsagris@uoc.gr>
Depends: R (≥ 4.0)
Imports: cluster, graphics, grDevices, MASS, rangen, Rfast, stats
Suggests: Compositional, CompositionalZADR, mziln, Rfast2
Description: Modelling structural zeros in compositional data using a conditional logistic normal model as described by Aitchison (1986), where MLE (Maximum Likelihood Estimation) is performed via the EM (Expectation-Maximization) algorithm. The relevant paper is Alzeley and Tsagris (2026) <doi:10.48550/arXiv.2608.29954>.
License: GPL-2 | GPL-3 [expanded from: GPL (≥ 2)]
NeedsCompilation: no
Packaged: 2026-10-06 20:14:19 UTC; mtsag
Repository: CRAN
Date/Publication: 2026-10-06 22:00:31 UTC

Modelling Compositional Data with Zero Values

Description

Modelling Compositional Data with Zero Values.

Details

Package: Compositionalcln
Type: Package
Version: 1.1
Date: 2026-10-06

Maintainers

Michail Tsagris <mtsagris@uoc.gr>.

Author(s)

Michail Tsagris mtsagris@uoc.gr

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data.


Maximum likelihood estimation of the conditional logistic normal model

Description

Maximum likelihood estimation of the conditional logistic normal model.

Usage

boot.clnmle(y, tol = 1e-6, maxit = 500, R = 1000)

Arguments

y

A numerical matrix with compositional data with zero values.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

R

The number of bootstrap resamples to generate.

Details

The function performs non-parametric bootstrap for the conditional multivariate logistic normal distribution (Alzeley and Tsagris, 2026).

Value

A list including:

m

The bootsrap mean vectors in the Euclidean space.

mesi

The bootstrap mean vectors in the simplex space.

S

The bootstrap estimated covariance matrices.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.reg

Examples

y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- boot.clnmle(y, R = 300)

Bootstrap for the conditional logistic normal regression

Description

Bootstrap for the conditional logistic normal regression.

Usage

boot.clnreg(y, x, tol = 1e-6, maxit = 500, R = 1000)

Arguments

y

A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated.

x

The predictor variable(s), they can be either continnuous or categorical or both.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

R

The number of bootstrap resamples to generate.

Details

Bootstrap to estimate the covariance matrix of the coefficients of the conditional logistic normal regression is performed.

Value

A list including:

beta

The bootstrapped beta coefficients.

sigma

The bootstrap estimate of the covariance matrix of the regression parameters.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.reg

Examples

x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.reg(y, x)

Non-parametric confidence region of the simplicial mean on a ternary diagram

Description

Non-parametric confidence region of the simplicial mean on a ternary diagram.

Usage

cln.bcr(y, tol = 1e-6, maxit = 500, alpha = 0.05, R = 1000,
dg = TRUE, hg = TRUE, ploty = TRUE)

Arguments

y

A matrix with the compositional data (with zero values).

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

alpha

The significance level to construct the (1-alpha)% confidence region.

R

The number of bootstrap resamples to generate.

dg

Do you want diagonal grid lines to appear? If yes, set this TRUE.

hg

Do you want horizontal grid lines to appear? If yes, set this TRUE.

ploty

Do you want to plot the data as well?

Details

A non-parametric confidence region of the mean in S^2 is constructed based on the method of Fiksel, Zeger and Datta (2022).

Value

A ternary diagram with the simplicial mean and the non-parametric confidence region.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Fiksel J., Zeger S. and Datta A. (2022). A transformation-free linear regression for compositional outcomes and predictors. Biometrics, 78(3): 974–987.

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.cr, cln.contour, cln.mle

Examples

y <- as.matrix(iris[1:50, 1:3])
y <- y / rowSums(y)
ind <- sample(50, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
cln.bcr(y, R = 300, ploty = FALSE)

Biplot based on the conditional logistic normal model

Description

Biplot based on the conditional logistic normal model.

Usage

cln.biplot(y, m, S)

Arguments

y

A vector or a matrix with compositional data with zero values.

m

The mean vector in R^{D-1}.

S

The covariance matrix in R^{D-1}.

Details

The function first transforms the covariance matrix that is based on the alr transformation to the covariance matrix based on the clr transformation. Then, for the vectors with zero values it uses the vectors with the estimated zis on the alr-space (see Alzeley and Tsagris, 2026) from the EM algorithm. Then it produces the classical biplot. The observations that contain zero values will be plotted with a red colour and the observations with no zero values are plotted with a grey colour.

Value

The biplot of the alr transformed data.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data.

See Also

cln.biplot, cln.contour, cln.mle

Examples

y <- as.matrix(iris[, 1:4])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(4, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
cln.biplot(y, mod$m, mod$S)

Many conditional logistic normal regressions given some covariates

Description

Many conditional logistic normal regressions given some covariates.

Usage

cln.condregs(y, X, prior, tol = 1e-6, maxit = 500)

Arguments

y

A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated.

X

A matrix with the predictor variable(s). A regression model will be fit to each of them, separately.

prior

A matrix or a vector with some predictors that will be in the model.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

Details

A conditional logistic normal regression is being fitted for each of the columns of X. The difference from cln.regs, is that this time, each of the the regression models fit contain the prior variables.

Value

A list including:

stat

A vector with the log-likelihood ratio test statistics.

pvalue

A vector with the corresponding logged p-values.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

See Also

cln.reg, cln.regs

Examples

x <- matrix( rnorm(150 * 10), ncol = 10 )
prior <- iris[, 4]
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.condregs(y, x, prior = prior)

Contour plot of the conditional logistic normal distribution in S^2

Description

Contour plot of the conditional logistic normal distribution in S^2.

Usage

cln.contour(m, S, n = 100, y = NULL, cont.line = FALSE)

Arguments

m

A value with the mean vector in S^2.

S

The covariance matrix in S^2.

n

The number of grid points to consider over which the density is calculated.

y

This is either NULL (no data) or contains a 3 column matrix with compositional data.

cont.line

Do you want the contour lines to appear? If yes, set this TRUE.

Details

The user can plot only the contour lines of a zero adjusted Dirichlet distribution with som given parameters, or can also add the relevant data should he/she wish to.

Value

A ternary diagram with the points and the conditional multivariate normal contour lines.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

dcln

Examples

y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)

cln.contour( m = mod$m, S = mod$S )

Confidence region of the simplicial mean on a ternary diagram

Description

Confidence region of the simplicial mean on a ternary diagram.

Usage

cln.cr(y, m, S, alpha = 0.05, dg = TRUE, hg = TRUE, ploty = TRUE)

Arguments

y

A matrix with the compositional data (with zero values).

m

A value with the mean vector in S^2.

S

The covariance matrix in S^2.

alpha

The significance level to construct the (1-alpha)% confidence region.

dg

Do you want diagonal grid lines to appear? If yes, set this TRUE.

hg

Do you want horizontal grid lines to appear? If yes, set this TRUE.

ploty

Do you want to plot the data as well?

Details

The confidence region of the mean in S^2 is constructed and then mapped back onto the simplex.

Value

A ternary diagram with the simplicial mean and the confidence region.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.bcr, cln.contour, cln.mle

Examples

y <- as.matrix(iris[1:50, 1:3])
y <- y / rowSums(y)
ind <- sample(50, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
m <- mod$m
S <- mod$S
cln.cr(y, m, S, ploty = FALSE)

James test for equality of two mean vectors based on the conditional logistic normal model

Description

James test for equality of two mean vectors based on the conditional logistic normal model.

Usage

cln.james(y1, y2, a = 0.05, tol = 1e-6, maxit = 500)

Arguments

y1

A numerical matrix with compositional data with zero values.

y2

A numerical matrix with compositional data with zero values.

a

The significance level, set to 0.05 by default.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

Details

The function fits the conditional multivariate logistic normal distribution (Alzeley and Tsagris, 2026) to each group, estimates the mean vectors and covariance matrices and uses them to perform the James test for equality of two mean vectors without assuming equality of the covariance matrices.

Value

A vector with 4 numbers, the value of the test statistic, the p-value, the correction and the corrected critical value.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.mle, cln.reg

Examples

y1 <- as.matrix(iris[1:75, 1:4])
ind <- sample(75, 10)
for ( k in ind )  y1[k, sample(3, 1)] <- 0
y1 <- y1 / rowSums(y1)

y2 <- as.matrix(iris[76:150, 1:4])
ind <- sample(75, 10)
for ( k in ind )  y2[k, sample(3, 1)] <- 0
y2 <- y2 / rowSums(y2)
cln.james(y1, y2)

Maximum likelihood estimation of the conditional logistic normal model

Description

Maximum likelihood estimation of the conditional logistic normal model.

Usage

cln.mle(y, tol = 1e-6, maxit = 500)

Arguments

y

A numerical matrix with compositional data with zero values.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

Details

The function fits the conditional multivariate logistic normal distribution (Alzeley and Tsagris, 2026).

Value

A list including:

m

The mean vector in the Euclidean space.

mesi

The mean vector in the simplex space.

S

The estimated covariance matrix.

loglik

The log-likelhiood value.

patterns

A matrix with the patterns of zeros, where the value of 0 indicates the presence of a zero, and the last column contains the percentage of occurrence each pattern. This is useful for random values simulation.

iters

The number of iterations required by the EM algorithm.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.reg

Examples

y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
cln.mle(y)

Principal component analysis based on the conditional logistic normal model

Description

Principal component analysis based on the conditional logistic normal model.

Usage

cln.pca(y, m, S)

Arguments

y

A vector or a matrix with compositional data with zero values.

m

The mean vector in R^{D-1}.

S

The covariance matrix in R^{D-1}.

Details

The function first transforms the covariance matrix that is based on the alr transformation to the covariance matrix based on the clr transformation. Then, for the vectors with zero values it uses the vectors with the estimated zis on the alr-space (see Alzeley and Tsagris, 2026) from the EM algorithm. Then it produces the classical principal component analysis.

Value

A list including:

values

A vector with the eigenvalues.

prop

A vector with the proportion of the eigenvalues

loadings

A matrix with the estimated loadings

scores

A vector with the estimated scores.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data.

See Also

cln.biplot, cln.contour, cln.mle

Examples

y <- as.matrix(iris[, 1:4])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
pc <- cln.pca(y, mod$m, mod$S)
sc <- pc$scores[, 1:2]  # centered, but unstandardised
# orthonormal basis of the clr plane of 3 parts (ilr basis)
Psi <- cbind( c(1, -1, 0) / sqrt(2), c(1,  1, -2) / sqrt(6) )
est <- tcrossprod( sc, Psi)
phat <- exp(est)  ;  phat <- est / rowSums(est)
ternary(phat)

Conditional logistic normal regression

Description

Conditional logistic normal regression.

Usage

cln.reg(y, x, tol = 1e-6, maxit = 500, seb = FALSE, xnew = NULL)

Arguments

y

A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated.

x

The predictor variable(s), they can be either continnuous or categorical or both.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

seb

If you want the standard errors of the regression coefficients set this to TRUE.

xnew

If you have new data use it, otherwise leave it NULL.

Details

The conditional logistic normal regression is being fitted. The likelihood conists of two components. The contributions of the non zero compositional values and the contributions of the compositional vectors with at least one zero value. The second component may have many different sub-categories, one for each pattern of zeros.

Value

A list including:

patterns

A matrix with the patterns of zeros, where the value of 0 indicates the presence of a zero, and the last column contains the percentage of occurrence each pattern. This is useful for random values simulation.

iters

The number of iterations required by the EM algorithm.

S

The covariance of the latent residuals.

beta

The beta coefficients.

loglik

The value of the log-likelihood.

est

The fitted or the predicted values (if xnew is not NULL).

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.mle

Examples

x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.reg(y, x)

Many conditional logistic normal regressions

Description

Many conditional logistic normal regressions.

Usage

cln.regs(y, X, tol = 1e-6, maxit = 500)

Arguments

y

A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated.

X

A matrix with the predictor variable(s). A regression model will be fit to each of them, separately.

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

Details

A conditional logistic normal regression is being fitted for each of the columns of X.

Value

A list including:

stat

A vector with the log-likelihood ratio test statistics.

pvalue

A vector with the corresponding logged p-values.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

See Also

cln.reg

Examples

x <- matrix( rnorm(150 * 10), ncol = 10 )
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.regs(y, x)

Residual diagnostic plot for the CLN and CLNR models

Description

Residual diagnostic plot for the CLN and CLNR models.

Usage

cln.residplot(y, m = NULL, S, x = NULL, beta = NULL)

Arguments

y

A numerical matrix with the compositional data, possibly containing zero values. The last column is the reference component of the additive log-ratio transformation.

m

The estimated mean vector of the alr transformed data, of length D-1, as returned by cln.mle. It is ignored when x is supplied.

S

The estimated (D-1) \times (D-1) covariance matrix of the (residuals of the) alr transformed data, as returned by cln.mle or cln.reg.

x

The predictor variable(s), the same as those given to cln.reg. Do not supply a design matrix, since the intercept is added internally. If NULL (default), the model without covariates is assumed.

beta

The p \times (D-1) matrix of the estimated regression coefficients, as returned by cln.reg, where the first row contains the intercepts. Required if x is supplied.

Details

For the i-th composition let b_i be the vector of the log-ratios among its C_i nonzero components, each taken against the last nonzero one. Under the model, b_i follows a multivariate normal distribution with mean Q_i \mu_i and covariance matrix Q_i \Sigma Q_i^\top, where Q_i depends only on the zero pattern and \mu_i is either the common mean or B^\top x_i. The squared Mahalanobis distance

m_i = (b_i - Q_i \mu_i)^\top (Q_i \Sigma Q_i^\top)^{-1} (b_i - Q_i \mu_i)

is compared to a chi-square distribution with c_i = C_i - 1 degrees of freedom. The chi-square probabilities are transformed to normal scores, z_i = \Phi^{-1}(G_{c_i}(m_i)), which have a common scale for all rows, and are plotted against the standard normal quantiles. Only the observed part of each vector is used, so no zero value is replaced.

The reference distribution ignores that the parameters were estimated from the same data, so for small samples or many predictors the plot may look slightly better than the fit deserves. Compositions with only one nonzero component contain no log-ratio and are omitted (NA).

Value

The QQ-plot is produced and a list including:

m2

The squared Mahalanobis distance m_i of the observed log-ratios of each row. NA for rows with a single nonzero component.

dof

The degrees of freedom of each distance, that is, the number of nonzero components minus one.

cdf

The value of the chi-square cumulative distribution function with dof degrees of freedom evaluated at m2. Under a correct model it is approximately uniform. It is not a p-value, whose upper tail is 1 - cdf.

z

The normal scores of cdf, that is, the quantities plotted. Under a correct model they are approximately standard normal, and large positive values indicate a poorly fitted composition.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (2003). The statistical analysis of compositional data. New Jersey: Reprinted by The Blackburn Press.

See Also

cln.mle, cln.reg

Examples

x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
## without covariates
mod <- cln.mle(y)
cln.residplot(y, m = mod$m, S = mod$S)

## with covariates
mod <- cln.reg(y, x)
cln.residplot(y, S = mod$S, x = x, beta = mod$beta)

Density of the conditional logistic normal model

Description

Density of the conditional logistic normal model.

Usage

dcln(y, m, S, logged = FALSE)

Arguments

y

A vector or a matrix with compositional data with zero values.

m

The mean vector in R^{D-1}.

S

The covariance matrix in R^{D-1}.

logged

If you want the log of the density set this TRUE, otherwise leave it FALSE.

Details

The function computes the density values of the conditional multivariate normal model.

Value

The density values at the given compositional data x.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Aitchison J. (1986). The statistical analysis of compositional data, pages 272–274.

See Also

cln.mle, cln.contour

Examples

x <- as.vector(iris[, 4])
y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
f <- dcln(y, mod$m, mod$S)

Forward-backward with early dropping variable selection

Description

Forward-backward with early dropping variable selection.

Usage

fbed.cln(y, X, prior = NULL, alpha = 0.05, univ = NULL, K = 0,
tol = 1e-6, maxit = 500)

Arguments

y

A matrix with the compositional data (dependent variable). The number of observations (vectors) with no zero values should be more than the columns of the predictor variables. Otherwise, the initial values will not be calculated.

X

A matrix with the available predictor variables.

prior

A matrix or a vector with some predictors that will be in the model.

alpha

The significance level.

univ

If you have the test statistics and the p-values (the output of cln.regs or of cln.condregs) of the simple regressions pass it here.

K

How many times should the FBED be repeated?

tol

The tolerance value to terminate the EM algorithm.

maxit

The maximum number of iterations allowed for the EM algorithm.

Details

The function performs the FBED variable selection algorithm of Borboudakis and Tsamardinos (2019).

Value

A list including:

univ

The result of the simple regressions, the same result as cln.regs. This is a list that contains the log-likelihood ratio test statistics and their corresponding logged p-values.

res

A matrix with the selected variables, their test statistic and their logged p-value.

info

A matrix with two columns, the number of variables selected at each step and the number of tests performed at each step. Each step means at each value of K.

runtime

The running time of the algorithm.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Alzeley O. and Tsagris M. (2026). Modelling compositional data with structural zero values https://arxiv.org/pdf/2608.29954

Borboudakis G. and Tsamardinos I. (2019). Forward-backward selection with early dropping. Journal of Machine Learning Research, 20(8): 1-39.

See Also

cln.reg

Examples

y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
X <- matrix(rnorm(150 * 30), ncol = 30)
mod <- fbed.cln(y, X)

Ternary diagram

Description

Ternary diagram.

Usage

ternary(y, dg = FALSE, hg = FALSE, m = NULL, colour = NULL)

Arguments

y

A matrix with the compositional data (with zero values present).

dg

Do you want diagonal grid lines to appear? If yes, set this TRUE.

hg

Do you want horizontal grid lines to appear? If yes, set this TRUE.

m

If you know the estimated simplicial mean pass it here, otherwise leave it NULL. It will appear with the symbol +.

colour

If you want the points to appear in different colour put a vector with the colour numbers or colours.

Details

There are two ways to create a ternary graph. We used here that one where each edge is equal to 1 and it is what Aitchison (1986) uses. For every given point, the sum of the distances from the edges is equal to 1. Horizontal and or diagonal grid lines, and the mean can appear. Zeros in the data appear with red x.

Value

The ternary plot. Additionally, horizontal or diagonal grid lines can appear as well.

Author(s)

Michail Tsagris.

R implementation and documentation: Michail Tsagris mtsagris@uoc.gr.

References

Aitchison, J. (1983). Principal component analysis of compositional data. Biometrika 70(1): 57–65.

Aitchison J. (1986). The statistical analysis of compositional data. Chapman & Hall.

See Also

cln.contour

Examples

y <- as.matrix(iris[, 1:3])
y <- y / rowSums(y)
ind <- sample(150, 10)
for ( k in ind )  y[k, sample(3, 1)] <- 0
y <- y / rowSums(y)
mod <- cln.mle(y)
ternary(y, hg = TRUE, dg = TRUE, m = mod$mesi)