hv-metacognition

Predict whether a training run will succeed from its first 10%. A ridge regressor on 19 features extracted from the early training trajectory predicts final validation accuracy with LOO R² = 0.90 and MAE = 0.066.

Requires PyTorch (for the base models) and NumPy (for the meta-regressor). Runtime: ~100 s for the full benchmark on CPU (48 training runs + LOO).

The finding

Two surprising results from the benchmark:

1. Gradient stability dominates. The single most important feature is log_grad_mean (mean log gradient norm over the observation window), not val_acc_end or log_loss_end:

feature standardized importance
log_grad_mean 0.172
log_loss_start 0.097
log_param_end 0.095
val_acc_start 0.087
param_growth 0.081
log_n_params 0.079
val_acc_slope 0.050
val_acc_end 0.041

If you want to know whether a training run will succeed, look at gradient health, not at early loss or accuracy.

2. Ten percent is as informative as seventy. The fraction sweep:

observed fraction R² MAE
0.10 0.900 0.066
0.15 0.913 0.066
0.20 0.908 0.067
0.30 0.933 0.056
0.40 0.939 0.053
0.50 0.929 0.058
0.70 0.952 0.047

The elbow is at 10%. Watching training for longer adds very little predictive power.

Headline numbers

model R² MAE RMSE
constant (mean of others) −0.043 0.232 0.280
single: val_acc_end 0.586 0.153 0.176
single: log_loss_end 0.443 0.182 0.204
single: val_acc_slope 0.527 0.166 0.188
full meta-regressor (30%) 0.933 0.056 0.071

Calibration on 48 held-out runs: mean |err| = 0.044, max |err| = 0.146, 45/48 within ±0.10.

13/13 consistency checks pass.

The 19 features

Extracted from the first observed_fraction of a training run:

group features
Loss level log_loss_start, log_loss_end, log_loss_mean, log_loss_std
Loss trend log_loss_slope, log_loss_curvature, log_loss_residual, log_loss_improvement
Gradient log_grad_mean, log_grad_slope, log_grad_std
Weights param_growth, log_param_end
Early val val_acc_start, val_acc_end, val_acc_slope

The meta-dataset

48 training runs:

  • Tasks: Fibonacci-mod-8, mod-16, mod-24
  • Seeds: 0 and 1 per configuration
  • Final accuracy range: 0.155 to 0.967, mean 0.644, std 0.274. The wide variance is what makes the meta-regression meaningful.

How to use

from hv_metacognition import (
    Task, ModelConfig, TASKS,
    train_with_trajectory, extract_features, features_to_vector,
    RidgeMetaRegressor, loo_evaluate,
)

# Train a small model, record trajectory
task = TASKS['fib16']
config = ModelConfig(d_model=32, n_heads=4, n_layers=2, activation='gelu')
traj = train_with_trajectory(task, config, steps=200, seed=0)

# Extract features from the first 30%
feats = extract_features(traj, observed_fraction=0.30)
x = features_to_vector(feats)
print(f"predicted final: some value, actual: {traj['final_acc']}")

# To train a meta-model, see the benchmark functions

---

Command line:

```bash
python hv_metacognition.py                  # full benchmark (~100s)
python hv_metacognition.py collect          # meta-dataset collection only
python hv_metacognition.py metacog          # meta-regression only
python hv_metacognition.py fraction         # fraction sweep
python hv_metacognition.py predictions      # per-run predictions

Intended use

  • Early-stopping guidance. Run 10% of the planned budget, extract features, predict final accuracy. Kill runs unlikely to succeed.
  • Hyperparameter sweeps. Rank candidate configurations by predicted final accuracy from short runs, then train only the top-k to completion.
  • Educational. The features are interpretable and the regressor is linear. The finding about log_grad_mean is a concrete lesson about training dynamics.

Limitations

  • Small models, short runs. d_model ≤ 64, 200 steps, 24-token sequences. The meta-features might not transfer to real language models with millions of parameters and billions of tokens.
  • 48 runs is a small meta-dataset. LOO cross-validation with 48 samples has high variance in the R² estimate. Treat the number as a point estimate, not a confidence interval.
  • Synthetic tasks only. Fibonacci-mod-P is deterministic. Real LM tasks have different dynamics.
  • Linear meta-regressor. A nonlinear meta-model might extract more signal. The linear ridge is a fair baseline for "how much is in the features alone."
  • Predicts within a run, not across tasks. This is not hyperparameter selection for a new task; it's prediction of the outcome of a run already in progress.
  • Requires the trajectory, not just the model. The features come from observed training dynamics. A newly-initialized model has no gradient-norm history yet.
  • No confidence intervals on predictions. The LOO residuals give an empirical error distribution but the meta-model doesn't emit a per-prediction CI.

Comparison to related work

tool scope predicts deps
learning-curve extrapolation (e.g., Domhan et al.) final loss PyTorch
Freeze-Thaw BOHB hyperparameter sweep various
hv-metacognition final accuracy from 10% trajectory PyTorch + numpy

Learning-curve extrapolation fits the loss trajectory and predicts future loss. Metacognition extends this by combining loss, gradient, weight, and accuracy signals into a single prediction, and by identifying which signal dominates. The answer — gradient norm — is not the one that curve-fitting literature would predict.

Files

  • hv_metacognition.py — full source (PyTorch + NumPy)
  • config.json — feature list, hyperparameters, benchmark results
  • example.py — usage demonstrations
  • README.md — this card

Citation

If you use this, cite it as hv-metacognition from the zeechimp HuggingFace organization.

Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support