Skip to contents

What the target changes

Let the training sample contain nn observations. Observation ii has binary response yi∈{0,1}y_i\in\{0,1\}. There are GG prespecified groups. After training-only transformations, zigz_{ig} is the retained predictor vector for group gg, and bgb_g is its coefficient vector in the same coordinates. The unpenalized intercept is b0b_0, and ηi=b0+∑g=1Gzig𝖳bg\eta_i=b_0+\sum_{g=1}^{G}z_{ig}^{\mathsf T}b_g is the linear predictor.

Conditionally on the training-derived groupwise Firth targets tgt_g, Equation (1) defines the binomial estimation objective: mean negative Bernoulli log-likelihood plus a mixed group-sparsity/target-shrinkage penalty.

1n∑i=1n{log⁡(1+eηi)−yiηi}+λ∑g=1Gwg{α‖bg‖2+1−α2‖bg−dtg‖22}. \frac{1}{n}\sum_{i=1}^n\{\log(1+e^{\eta_i})-y_i\eta_i\} +\lambda\sum_{g=1}^G w_g \left\{\alpha\lVert b_g\rVert_2+ \frac{1-\alpha}{2}\lVert b_g-dt_g\rVert_2^2\right\}. \tag{1}

In Equation (1), λ>0\lambda>0 is finite, α,d∈[0,1]\alpha,d\in[0,1], ‖⋅‖2\lVert\cdot\rVert_2 is the Euclidean norm, and wgw_g is the square root of the group’s retained rank. Targets and transformations are estimated from the current training partition, not validation or test responses. The targets are training-data dependent; this conditional formulation does not assert an unconditional risk guarantee.

At d=0d=0, Equation (1) reduces to zero-target group elastic net. At α=1\alpha=1, the quadratic term vanishes and dd has no effect. Positive dd shifts the shrinkage destination; it does not guarantee better prediction or exact group-support recovery. The intercept is not part of the penalty. See the package’s sglasso-binomial reference and the logistic research code for the study-specific implementation.

A small fit

set.seed(19)
x <- matrix(rnorm(240), 60, 4)
y <- rbinom(60, 1, plogis(x[, 1]))
group <- c(1, 1, 2, 2)
fit <- sglasso(x, y, group, family="binomial",
               lambda=c(0.3, 0.1), d=c(0, 0.5))
predict(fit, newx=x[1:3, ], lambda=0.1, d=0.5, type="response")
#> [1] 0.2800003 0.5824570 0.5030327

These are probabilities for y = 1, not automatically fitted class labels. For classification, declare a threshold using training/validation information or an application-specific decision policy. Never optimize it on the test set. The first three training rows above illustrate syntax only.

Training-only targets in CV

Binomial CV recomputes standardization, group orthonormalization and Firth targets inside each training fold. It minimizes observation-weighted out-of-fold log-loss. Alpha remains fixed within a call.

set.seed(29)
cv <- cv.sglasso(x, y, group, family="binomial",
                 lambda=c(0.3, 0.1), d=c(0, 0.5),
                 alpha=0.5, nfolds=3)
#> Warning: CV selected a finite lambda-grid endpoint; performance beyond the grid
#> is unassessed.
probability <- predict(cv, newx=x[1:3, ], s="opt", type="response")
probability
#> [1] 0.2800003 0.5824570 0.5030327

The tiny grid is for documentation testing, not a recommendation for a study. Endpoint warnings are informative: a better value beyond the requested grid has not been ruled out. A supplied lambda vector is absolute; automatic paths are training-fold-specific relative paths aligned by grid position. No universal positive-target null-model lambda maximum is claimed.

Numerical scope

The compiled solver uses IRLS/proximal-Newton updates, profiled block updates and monotone accelerated proximal-gradient fallback. Required Firth targets must pass their adjusted-score/Fisher-step checks; they are not replaced by ordinary MLE or ridge on failure. Requested finite path points must pass their numerical gates. No guarantee covers every possible input.

Missing or nonnumeric predictors are rejected, not imputed. Constant columns and rank-deficient groups are handled by the documented training-only transformations. Gaussian screening and stopping options are not simply transplanted into the binomial interface. Read the reference before changing eps, max_iter or binomial.control.

The across-fold cvse is not an independent-test confidence interval. ROC-AUC measures ranking; log-loss and Brier score assess probabilities; MCC assesses classification under a declared threshold. Good performance on one is not automatically good performance on another.