---
title: "Ensemble estimators and residual offsets"
output: rmarkdown::html_vignette
bibliography: ../inst/REFERENCES.bib
link-citations: true
vignette: >
  %\VignetteIndexEntry{Ensemble estimators and residual offsets}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(ssel)
```

This article explains the estimators used by the supervised helper family.
It is important to keep them separate: `trainRegressionModel()` constructs
cell-level OOF and unseen-row ensembles, `oofEnsemble()` independently
reconstructs training-row residuals, `auditQuantiles()` summarizes those
residuals as signed offsets, and the two delivery helpers apply different
selection rules. `modelPipeline()` orchestrates these helpers; it is not an
additional statistical estimator.

## Choose the right API layer

The two training helpers are active at different levels; they are not duplicate
implementations.

| Task | Public helper |
|---|---|
| run the complete supervised workflow from assembled dataset CSVs | `modelPipeline()` |
| discover all TRAIN splits and fit them with ssel's standard defaults | `trainModel()` |
| fit an explicit cell grid with custom resampling or tuning controls | `trainRegressionModel(TRAIN = TRUE)` |
| evaluate fitted cells, select deliveries, and assemble user-facing products | `predictModel()` |
| obtain raw cell metrics and long predictions without delivery selection | `trainRegressionModel(TRAIN = FALSE)` |

`trainModel()` calls the engine in training mode and returns after model files
are written. The metrics and response paths in its signature are prepared for
the later evaluation phase; the wrapper does not populate them. Evaluation mode
is likewise more than `predict()`: it first rebuilds OOF scores and ensembles,
then predicts unseen rows. Most applications should therefore enter through
`modelPipeline()` or the phase wrappers and use the cell engine directly only
for controls those wrappers do not expose.

## Notation and cells

Let $r$ identify a response, $d$ a dataset, $m$ a learner, and $i$ a
training row. A **cell** is one response--dataset pair $(r,d)$. For a cell,
$y_i$ is the observed response and $\hat y_{mi}^{OOF}$ is learner $m$'s
out-of-fold (OOF) prediction. RMSE and response predictions have the response's
units; squared sample correlation $R^2$ and normalized weights are
dimensionless.

The principal API mapping is:

| Mathematical role | Public helper |
|---|---|
| discover the standard fitting grid | `trainModel()` |
| fit explicit cells and retain selected-tuning OOF rows | `trainRegressionModel(TRAIN = TRUE)` |
| construct cell-level OOF and unseen-row ensembles | `trainRegressionModel()` |
| independently reconstruct training-row OOF residuals | `oofEnsemble()` |
| summarize residuals as signed offsets | `auditQuantiles()` |
| select a response-level cell | `predictModel()` |
| select final row-level RMSE candidates | `aggregateResponses()` |
| compute the package optimism diagnostic | `auditOverfit()` |

## Fitting and OOF scores

Learners are fitted separately with `caret::train()` [@kuhn2008]; current
resampling, tuning, and reproducibility controls are documented in the
[caret training and tuning guide](https://topepo.github.io/caret/model-training-and-tuning.html).
Through `trainModel()`, ssel requests five-fold cross-validation, RMSE tuning,
selected-tuning holdout predictions, and `tuneLength = 10`. A direct
`trainRegressionModel()` call instead uses its `.numberCV`, `.tuneLength`, and
`.metric` arguments. The tuning length requests method-specific grid
granularity; it is not a number of points per tuning parameter.

Separate learner calls do not receive a common explicit resampling index.
Their saved OOF rows therefore support learner-specific score estimation, but
ssel does not claim identical folds across learners.

Within one cell, the current evaluation engine uses the first requested learner
that reaches scoring to establish the shared ordered observation vector. Later
learners are scored against that vector; their own saved observations and row
indices are not checked for equality. The equations below therefore describe
the intended consistent-fit case. If fitted objects carry inconsistent OOF
rows, learner order becomes material and the cell should be treated as invalid
input rather than as a comparable ensemble.

For each retained learner in cell $(r,d)$, evaluation reports

$$
\operatorname{RMSE}_{mrd}
=\left\{\frac{1}{n}\sum_i(y_i-\hat y_{mi}^{OOF})^2\right\}^{1/2}.
$$

$$
R^2_{mrd}=\operatorname{cor}(y,\hat y_m^{OOF})^2.
$$

The package defines two cell-local weighting policies:

$$
w^{R2}_{mrd}=\frac{R^2_{mrd}}{\sum_jR^2_{jrd}},
\qquad
w^{RMSE}_{mrd}
=\frac{\operatorname{RMSE}_{mrd}^{-1}}
       {\sum_j\operatorname{RMSE}_{jrd}^{-1}}.
$$

The corresponding OOF ensemble is

$$
\hat y_{ird}^{g,OOF}=\sum_m w^g_{mrd}\hat y_{mi}^{OOF},
\qquad g\in\{R2,RMSE\}.
$$

These normalized scores are package policies, not estimates of theoretically
optimal ensemble weights and not guarantees that an ensemble improves every
learner.

## Unseen-row ensembles are cell-local

During OOF evaluation, `trainRegressionModel()` stores the exact named
$R^2$ and RMSE lists for every cell. Before predicting unseen rows in
$(r,d)$, it restores that cell's lists and intersects their names with the
learners that produced valid predictions for that same cell. A missing model
or a missing prediction removes a learner only from the current cell.

The unseen-row pass uses the same normalization formulas. Its documented
fallback substitutes equal weights when the retained $R^2$ values, or the
reciprocal RMSE values, contain a missing value or sum to zero. A cell without
stored OOF score state has no unseen-row ensemble. Thus one cell neither
borrows scores from nor prunes methods for another cell.

## Independent OOF residual reconstruction

`oofEnsemble()` does not reuse the preceding score state. It recomputes a
method-level RMSE from each fitted object's saved, projected OOF predictions
and defines independent inverse-RMSE weights $\tilde w_{mrd}$. Its cast key is
jointly $h=(i,o)$, where $i$ is the caret row index and $o$ the averaged
saved observation. With $A_h$ denoting learners available at that key,

$$
\hat y_h^{OOF}=\Pi_{[a,b]}\!\left(
\frac{\sum_{m\in A_h}\tilde w_{mrd}\tilde y_{mh}}
     {\sum_{m\in A_h}\tilde w_{mrd}}
\right).
$$

$$
e_h=o-\hat y_h^{OOF}.
$$

Available learners are renormalized row by row. If saved observations differ
between learners for one row index, they form different joint keys; repeated
row indices and sample identifiers can therefore be legitimate output rows.

For valid bounds $a\le b$, the interval projection used above is

$$
\Pi_{[a,b]}(z)=\min\{b,\max\{a,z\}\}.
$$

This is the ordinary projection onto a closed interval [@parikh2014]. Bounds
are not proactively reordered, so this name applies only when $a\le b$.

## Signed residual offsets

`auditQuantiles()` pools reconstructed residual rows separately within each
cell. For $p\in\{0.50,0.90,0.95\}$, sort the $n$ residuals and define

$$
H=(n-1)p+1.
$$

$$
j=\lfloor H\rfloor,\qquad \gamma=H-j.
$$

R's [type-7 sample quantile](https://www.stat.ethz.ch/R-manual/R-devel/library/stats/html/quantile.html)
is then

$$
q_{p,rd}=(1-\gamma)e_{(j)}+\gamma e_{(j+1)},
$$

with the endpoint convention described by @hyndman1996. The values are signed,
cell-pooled, and expressed in response units. Delivery helpers form

$$
Q_{p,ird}=\Pi_{[a,b]}\{\hat y_{ird}+q_{p,rd}\}.
$$

These are descriptive point-plus-residual-offset estimates. They are not
conditional predictive quantiles or intervals with a calibrated coverage
guarantee. Coverage-guaranteed procedures such as jackknife+ require fold- or
leave-one-out-specific predictions at each test point, which this estimator
does not construct [@barber2021].

## Two distinct selection rules

`predictModel()` first chooses one eligible cell per response. For
`ensemble.RMSE` it takes the first minimum-RMSE row; for `ensemble.R2` it takes
the first maximum $R^2$ row:

$$
d_r^{RMSE}=\operatorname*{first\,argmin}_{d}\operatorname{RMSE}_{rd}.
$$

$$
d_r^{R2}=\operatorname*{first\,argmax}_{d}R^2_{rd}.
$$

`aggregateResponses()` applies a later and different rule. Its candidates are
only unseen `ensemble.RMSE` points and reconstructed training OOF points. For
each sample $i$ and response $r$, it selects

$$
d^*_{ir}=\operatorname*{first\,argmax}_{d}R^2_{rd}.
$$

The selected candidate supplies all three signed offsets. Because the second
rule is grouped by sample and excludes `ensemble.R2` candidates, the two
selectors must not be interpreted as one global best-dataset operation.

## Optimism diagnostic

`auditOverfit()` compares projected in-sample RMSE with the arithmetic mean of
the fitted caret object's stored resample RMSE values:

$$
\operatorname{rel\_optimism}
=\frac{\overline{\operatorname{RMSE}}_{CV}-\operatorname{RMSE}_{in}}
       {\max\{\overline{\operatorname{RMSE}}_{CV},10^{-9}\}}.
$$

Its labels and thresholds are ssel policy. The value is a diagnostic, not a
hypothesis test, a universal measure of overfitting, or an input to final
aggregation.

## Executable residual-offset example

The example below creates one fitted-object fixture with two equal signed OOF
residuals. All three type-7 offsets must therefore equal one.

```{r residual-offset-example}
root <- tempfile("ssel-offsets-")
model_dir <- file.path(root, "model")
train_dir <- file.path(root, "train")
summary_dir <- file.path(root, "summary")
dir.create(model_dir, recursive = TRUE)
dir.create(train_dir)

data.table::fwrite(
  data.table::data.table(y = c(1, 3)),
  file.path(train_dir, "target_demo_TRAIN.csv")
)
data.table::fwrite(
  data.table::data.table(SampleID = c("S1", "S2")),
  file.path(train_dir, "target_demo_TRAIN_ids.csv")
)
fit <- list(
  pred = data.frame(
    tune = 1L, pred = c(0, 2), obs = c(1, 3), rowIndex = 1:2
  ),
  bestTune = data.frame(tune = 1L)
)
saveRDS(fit, file.path(model_dir, "lm_target_demo.Rds"))

offsets <- auditQuantiles(
  .path.model = model_dir,
  .path.train = train_dir,
  .path.summary = summary_dir,
  .response_pattern = "target",
  verbose = FALSE
)
offsets$quantiles
stopifnot(all(offsets$quantiles[, c("q_50", "q_90", "q_95")] == 1))

unlink(root, recursive = TRUE)
```

The complete argument, file, projection, skip, and side-effect contracts live
in the linked function reference pages. The iterative multi-response and
pseudo-label policies are described in their own articles.

## References
