TimeDiT rebuild (small)

Reimplementation of TimeDiT (Cao et al., arXiv 2409.02322) trained from scratch on a small subset of the Chronos/LOTSA corpus (Salesforce/lotsa_data), with the standard benchmark families (weather, traffic, electricity, solar, taxi, ETT) excluded from pre-training, so the ETTh1 evaluation is zero-shot.

Architecture: value tokens (no patching), two streams (noised target + clean condition) with AdaLN-zero conditioning, diffusion timestep embedding in the target tokens, DDPM cosine T=100 eps-prediction, DDIM sampling, unified mask mixture (random/block/stride/reconstruction).

Training

{
  "d": 384,
  "n_layers": 8,
  "heads": 6,
  "T": 100
}

Final training loss ~0.188 at 13000 steps (AdamW lr 1e-4 cosine, bf16, A10G, ~147 min). ~23M parameters.

ETTh1 zero-shot evaluation

Standard 12/4/4-month split; context 288; point forecast = mean of 10 DDIM samples (20 steps, bf16); MSE/MAE on train-split-normalized data; CRPS_sum = empirical CRPS of the 10-sample ensemble on channel-summed series; persistence_mse = naive last-value baseline on the same windows (lower MSE than the model at every horizon).

{
  "forecast": {
    "96": {
      "mse": 1.9348,
      "mae": 1.1477,
      "crps_sum": 0.3935,
      "persistence_mse": 1.2315,
      "windows": 43,
      "eval_seconds": 103.4
    },
    "192": {
      "mse": 1.8411,
      "mae": 1.0863,
      "crps_sum": 0.3496,
      "persistence_mse": 1.2166,
      "windows": 42,
      "eval_seconds": 143.0
    },
    "336": {
      "mse": 1.9515,
      "mae": 1.1303,
      "crps_sum": 0.3655,
      "persistence_mse": 1.2005,
      "windows": 39,
      "eval_seconds": 199.3
    }
  },
  "imputation": {
    "0.125": {
      "mse": 2.1559,
      "mae": 1.173
    },
    "0.25": {
      "mse": 1.8686,
      "mae": 1.0802
    },
    "0.5": {
      "mse": 1.9587,
      "mae": 1.137
    }
  }
}

Honest comparison with the paper

The paper's TimeDiT (pre-trained on ~27B observations, much larger model) reports ETTh1 forecast avg MSE 0.356 and imputation avg MSE 0.036. This 23M rebuild on a ~300MB LOTSA subset scores far worse on forecasting (worse than a persistence baseline) -- the gap is scale (data + model size), not a broken pipeline: the model trains stably and produces calibrated ensembles.

Downloads last month
56
Safetensors
Model size
22.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support