Training-Only Cross-Validation for Lambert Regression
cv.lambert.RdChoose a relative penalty-path position by the arithmetic mean of validation-fold squared errors, with fixed Lambert shape and fold-specific preprocessing.
Usage
cv.lambert(x, y, nfolds = 5L, foldid = NULL,
lambda_fraction = 10^seq(0, -3, length.out = 60L),
max_sweeps = 10000L, kkt_tol = 1e-7)Arguments
- x, y
Training predictors and response, as in
lambert.- nfolds
Number of folds, default five. Each training subset must contain at least three observations. Inferred from
foldidwhen that argument is supplied andnfoldsis omitted.- foldid
Optional integer fold label for every row, with consecutive labels starting at one and no empty folds. Store and reuse this vector for matched comparisons. Without it, balanced labels are randomly permuted using the caller's R random-number state. Use
set.seed()for reproducibility.- lambda_fraction
Strictly decreasing fractions in \((0,1]\); default 60 logarithmically spaced positions from 1 to 0.001.
- max_sweeps, kkt_tol
Solver controls described in
lambert.
Details
Each training fold estimates its own predictor centers, RMS scales, response mean and entry score. Candidate \(k\) in fold \(f\) uses \(\lambda_{k,f}=\lambda_{\max,f}\tau_k\), where \(\tau_k\) is the supplied fraction. The candidate is aligned by fraction, not by a common absolute lambda. The validation observations are never used to estimate these values.
For validation sets \(I_f\) and training-fold predictions \(\hat y_i\),
the score is
$$\mathrm{CV}(k)=\frac{1}{K}\sum_{f=1}^K
\frac{1}{|I_f|}\sum_{i\in I_f}(y_i-\hat y_i)^2.$$
This is an unweighted mean of the fold MSEs, including when fold sizes differ.
Only candidates with finite loss in every fold are eligible. Ties within
\(10^{-10}(1+\mathrm{CV}_{\min})\) favor the larger penalty fraction.
lambda.min is the selected fraction multiplied by the full training
entry score. A full-training path is fitted at every supplied fraction. Later positions
provide variable counts for plotting and do not affect CV selection. The
selected full and fold fits receive the reference active-curvature check.
A rejected selected fit is reported explicitly; another candidate is not
silently substituted.
An incomplete search is flagged by partial_search; it does not imply
that the selected fit failed. Solver convergence and curvature status must
be inspected separately. No one-standard-error rule or optional early
pruning is implemented in this initial package. CV error is a tuning score,
not an unbiased external performance estimate; use independent test data or
outer cross-validation to evaluate prediction. Changing the number of folds
or numerical controls is supported but departs from the study defaults.
Value
A list of class cv.lambert. Write \(L\) for the
number of fractions, \(K\) for the number of folds, and \(n\) for the
number of input rows. Components are:
- cvm
Length-\(L\) arithmetic means of fold MSEs. An ineligible candidate has value
Inf.- fold_loss
Numeric \(L\times K\) validation-MSE matrix, with
NAwhere a fold candidate is unavailable.- foldid, nfolds
Length-\(n\) fold labels and fold count.
- index
Selected path index, or
NA_integer_when none is eligible.- lambda_fraction, lambda
Length-\(L\) fractions and corresponding full-training penalty levels. These absolute levels need not equal fold levels.
- fraction.min, lambda.min
Selected fraction and full-training penalty;
NA_real_if no candidate is eligible.- lambert.fit
Full-training
lambertpath, including when no CV candidate is eligible. Its fields are documented inlambert.- nzero, support_tol
Length-\(L\) selected-variable counts and their threshold. Counts exclude the intercept and use original-scale coefficient magnitudes greater than \(10^{-8}\). Unverified counts are
NA.- full_path_complete
Whether every full-training position is verified; this is separate from fold-search completeness and selected-fit validity.
- ok, status
Logical selected-fit acceptance and character status.
no_valid_cvmeans no candidate had finite loss in all folds. Other status values followlambert. A negative-curvature selected fold can setstatus = "negative_curvature".- partial_search
Whether any candidate has a nonfinite CV score. A partial search can still have
ok = TRUE.- cv_diagnostics
Fold-stacked solver diagnostics, with additional
foldand validationlosscolumns.- attempts
Fold-stacked records for both starts, with a
foldcolumn. Theexecutedandreusedflags are described inlambert. Full-training attempts are inlambert.fit$attempts.- selected_curvature
Length-\(K\) list of selected-fold Hessian assessments;
NULLif no CV candidate is eligible. Full-training assessments are inlambert.fit$curvature.- preprocessing
Length-\(K+1\) list: each training fold, then the full training sample. Each entry contains
mx,my,sxandlambda0, the predictor means, response mean, RMS scales and entry score.- control, shape, call
Numerical settings, fixed shape (one), and call.
Completing the full-training path can cost more than fitting only its selected prefix. Plotting reads stored values and never fits models.
Examples
set.seed(12)
x <- matrix(rnorm(60 * 6), 60, 6)
y <- 2 * x[, 1] - x[, 2] + rnorm(60)
foldid <- sample(rep(1:5, length.out = 60))
fit <- cv.lambert(x, y, foldid = foldid,
lambda_fraction = c(1, 0.5, 0.2))
print(fit)
#> Cross-validated Lambert regression (5 folds)
#> Selected index: 3; full-training lambda: 0.3597978
#> Status: verified_stationary; partial search: FALSE
coef(fit)
#> s3
#> (Intercept) 0.1123142
#> V1 1.9992149
#> V2 -1.1984651
#> V3 0.0000000
#> V4 0.0000000
#> V5 0.0000000
#> V6 0.0000000
predict(fit, x[1:3, , drop = FALSE])
#> s3
#> [1,] -3.334891
#> [2,] 2.073636
#> [3,] -2.826032