hv-metacognition
Predict whether a training run will succeed from its first 10%. A ridge regressor on 19 features extracted from the early training trajectory predicts final validation accuracy with LOO R² = 0.90 and MAE = 0.066.
Requires PyTorch (for the base models) and NumPy (for the meta-regressor). Runtime: ~100 s for the full benchmark on CPU (48 training runs + LOO).
The finding
Two surprising results from the benchmark:
1. Gradient stability dominates. The single most important feature
is log_grad_mean (mean log gradient norm over the observation window),
not val_acc_end or log_loss_end:
| feature | standardized importance |
|---|---|
log_grad_mean |
0.172 |
log_loss_start |
0.097 |
log_param_end |
0.095 |
val_acc_start |
0.087 |
param_growth |
0.081 |
log_n_params |
0.079 |
val_acc_slope |
0.050 |
val_acc_end |
0.041 |
If you want to know whether a training run will succeed, look at gradient health, not at early loss or accuracy.
2. Ten percent is as informative as seventy. The fraction sweep:
| observed fraction | R² | MAE |
|---|---|---|
| 0.10 | 0.900 | 0.066 |
| 0.15 | 0.913 | 0.066 |
| 0.20 | 0.908 | 0.067 |
| 0.30 | 0.933 | 0.056 |
| 0.40 | 0.939 | 0.053 |
| 0.50 | 0.929 | 0.058 |
| 0.70 | 0.952 | 0.047 |
The elbow is at 10%. Watching training for longer adds very little predictive power.
Headline numbers
| model | R² | MAE | RMSE |
|---|---|---|---|
| constant (mean of others) | −0.043 | 0.232 | 0.280 |
single: val_acc_end |
0.586 | 0.153 | 0.176 |
single: log_loss_end |
0.443 | 0.182 | 0.204 |
single: val_acc_slope |
0.527 | 0.166 | 0.188 |
| full meta-regressor (30%) | 0.933 | 0.056 | 0.071 |
Calibration on 48 held-out runs: mean |err| = 0.044, max |err| = 0.146, 45/48 within ±0.10.
13/13 consistency checks pass.
The 19 features
Extracted from the first observed_fraction of a training run:
| group | features |
|---|---|
| Loss level | log_loss_start, log_loss_end, log_loss_mean, log_loss_std |
| Loss trend | log_loss_slope, log_loss_curvature, log_loss_residual, log_loss_improvement |
| Gradient | log_grad_mean, log_grad_slope, log_grad_std |
| Weights | param_growth, log_param_end |
| Early val | val_acc_start, val_acc_end, val_acc_slope |
The meta-dataset
48 training runs:
- Tasks: Fibonacci-mod-8, mod-16, mod-24
- Seeds: 0 and 1 per configuration
- Final accuracy range: 0.155 to 0.967, mean 0.644, std 0.274. The wide variance is what makes the meta-regression meaningful.
How to use
from hv_metacognition import (
Task, ModelConfig, TASKS,
train_with_trajectory, extract_features, features_to_vector,
RidgeMetaRegressor, loo_evaluate,
)
# Train a small model, record trajectory
task = TASKS['fib16']
config = ModelConfig(d_model=32, n_heads=4, n_layers=2, activation='gelu')
traj = train_with_trajectory(task, config, steps=200, seed=0)
# Extract features from the first 30%
feats = extract_features(traj, observed_fraction=0.30)
x = features_to_vector(feats)
print(f"predicted final: some value, actual: {traj['final_acc']}")
# To train a meta-model, see the benchmark functions
---
Command line:
```bash
python hv_metacognition.py # full benchmark (~100s)
python hv_metacognition.py collect # meta-dataset collection only
python hv_metacognition.py metacog # meta-regression only
python hv_metacognition.py fraction # fraction sweep
python hv_metacognition.py predictions # per-run predictions
Intended use
- Early-stopping guidance. Run 10% of the planned budget, extract features, predict final accuracy. Kill runs unlikely to succeed.
- Hyperparameter sweeps. Rank candidate configurations by predicted final accuracy from short runs, then train only the top-k to completion.
- Educational. The features are interpretable and the regressor is
linear. The finding about
log_grad_meanis a concrete lesson about training dynamics.
Limitations
- Small models, short runs. d_model ≤ 64, 200 steps, 24-token sequences. The meta-features might not transfer to real language models with millions of parameters and billions of tokens.
- 48 runs is a small meta-dataset. LOO cross-validation with 48 samples has high variance in the R² estimate. Treat the number as a point estimate, not a confidence interval.
- Synthetic tasks only. Fibonacci-mod-P is deterministic. Real LM tasks have different dynamics.
- Linear meta-regressor. A nonlinear meta-model might extract more signal. The linear ridge is a fair baseline for "how much is in the features alone."
- Predicts within a run, not across tasks. This is not hyperparameter selection for a new task; it's prediction of the outcome of a run already in progress.
- Requires the trajectory, not just the model. The features come from observed training dynamics. A newly-initialized model has no gradient-norm history yet.
- No confidence intervals on predictions. The LOO residuals give an empirical error distribution but the meta-model doesn't emit a per-prediction CI.
Comparison to related work
| tool | scope | predicts | deps |
|---|---|---|---|
| learning-curve extrapolation (e.g., Domhan et al.) | final loss | PyTorch | |
| Freeze-Thaw BOHB | hyperparameter sweep | various | |
| hv-metacognition | final accuracy from 10% trajectory | PyTorch + numpy |
Learning-curve extrapolation fits the loss trajectory and predicts future loss. Metacognition extends this by combining loss, gradient, weight, and accuracy signals into a single prediction, and by identifying which signal dominates. The answer — gradient norm — is not the one that curve-fitting literature would predict.
Files
hv_metacognition.py— full source (PyTorch + NumPy)config.json— feature list, hyperparameters, benchmark resultsexample.py— usage demonstrationsREADME.md— this card
Citation
If you use this, cite it as hv-metacognition from the zeechimp
HuggingFace organization.
- Downloads last month
- 18