The binomial API uses the accelerated IRLS/proximal-Newton solver with profiled block updates and monotone accelerated proximal-gradient fallback. The compiled core is a snapshot of the validated research solver, not the older Gaussian coordinate-descent implementation.
Details
On training-standardized, groupwise orthonormalized predictors, the objective is mean negative Bernoulli log-likelihood plus $$\lambda\sum_g w_g\{\alpha\|b_g\|_2 + (1-\alpha)\|b_g-d t_g\|_2^2/2\}.$$ Here \(b_g\) is the solver-coordinate coefficient vector for group \(g\), \(w_g\) is the square root of its retained rank, \(t_g\) is its training-only Firth slope estimate, \(d\in[0,1]\) scales the target, and \(\alpha\in[0,1]\) mixes group sparsity and quadratic shrinkage. The intercept is unpenalized. At \(d=0\) this is group elastic net; at \(\alpha=1\), the target has no effect. Firth fitting is skipped whenever the target cannot affect the objective. There is no fallback to MLE or ridge if a required Firth fit fails.
Constant columns (RMS centered scale at most \(10^{-8}\)) are removed internally and returned with zero coefficients. Each standardized group is reduced by SVD using relative tolerance \(10^{-10}\) times its largest dimension. The preprocessing record reports retained ranks, transformations, and excluded columns. Coefficients and predictions are on the original scale. Group labels need not be consecutive or numerically ordered. Missing data and nonnumeric predictors are rejected; no imputation is performed.
Automatic finite lambda grids start at the maximum weighted group null-score norm divided by alpha when alpha is positive. For alpha zero, the undivided score is only a reference scale. Neither reference is asserted to produce a null model for a positive target. The default path ends at 0.005 times its start. Specify a broader positive finite lambda grid when appropriate. Zero and infinite lambda are not part of this API. CV reports and warns about endpoint selection; it does not claim optimality outside the requested grid.
Binomial fits require standardize=TRUE, bilevel=FALSE,
beta_start=NULL, and eager coefficient transformation. Gaussian
SSR/SSR_fast rules and dfmax/gmax stopping are not transplanted to logistic
regression. The default screen setting selects the hybrid active set, with
full final KKT checks. SSR_fast is rejected. No path truncation by
training deviance is enabled. All requested points must pass numerical checks.
binomial.control accepts block_max_sweeps (4000),
block_chunk_sweeps (100), block_stall_window (200),
block_stall_relative_improvement (0.01),
apg_max_iterations (10000), apg_kkt_check_interval (5),
kkt_tolerance (eps), update_tolerance (1e-10),
intercept_tolerance (1e-12), max_intercept_iterations (100),
irls_max_outer (50), and irls_max_inner (2000).
When omitted, binomial eps is 1e-10 (Gaussian default unchanged).
max_iter caps each main iteration budget; it is not a wall-time limit.
Training Firth fits use BFGS with at most four 5000-iteration passes and five
score-polishing steps, requiring adjusted-score and Fisher-step norms <=1e-6.
cv.sglasso returns fold-local diagnostics and an observation-weighted
mean log-loss. Its cvse describes across-fold variation and is not an
independent-test confidence interval. Automatic folds are stratified and use
the caller's RNG; set a seed or supply fold for reproducibility.
Supplied lambda values are a fixed absolute grid; automatic CV paths are
aligned by relative position and recomputed within each training fold.
Selection uses minimum log-loss, breaking exact ties by path order (larger
lambda, then smaller d). A final full-training fit supplies predictions.
Effective degrees of freedom are not implemented: logLik returns
unpenalized log-likelihood with NA degrees of freedom and select
rejects binomial information-criterion selection. Use log-loss CV instead.