GON: Generalizable Outcome Networks

Code for GON: Generalizable Outcome Networks for Learning Action Effects in Physical Systems (anonymous submission).

GON learns how HVAC action sequences change a building's future from matched counterfactual branches: simulations that start from the same building state under the same weather and occupancy but apply different plans. An outcome model trained with a matched-branch loss predicts temperature and energy for a proposed plan; a conditional flow-matching generator proposes many plans for a requested temperature change; at deployment the outcome model scores the candidates and the first action of the best plan is executed, every 5 minutes. One frozen controller is evaluated zero-shot on 100 unseen Building2Building buildings.

Layout

gon/
  constants.py    observation/action layout, normalization, delivered-heat actuation map
  env.py          Building2Building env with the 5-minute step fix, episode loop, metrics
  baselines.py    tuned RBC (normalization anchor, logging policy) and PI controller
  branching.py    logged episodes, plan family, matched counterfactual branching (replay + diverge)
  data.py         history windows, matched-pair dataset, flow-generator items
  models.py       history encoder, outcome model + matched-branch loss, flow generator, checkpoints
  controller.py   the GON deployment controller
  utils.py        paths, splits, process pool
scripts/          collect_logs -> collect_branches -> train_encoder -> train_outcome -> train_flow -> evaluate

Setup

pip install -r requirements.txt
pip install -e <path-to-building2building>                 # EnergyPlus-based benchmark
export PYTHONPATH=$PWD:<path-to-building2building>:$PYTHONPATH   # `baselines` (tuned RBC) lives in that repository

Main settings

component setting
environment Building2Building OfficeSmall, winter (Jan 1 - Mar 31), 5-min step (minimum_system_timestep = 0), 25 920 steps
tasks task_const_e0 (comfort-focused, w_E = 0), task_const_e05 (energy-focused, w_E = 0.5), setpoint 21 degC
encoder causal transformer, 4 layers, 4 heads, width 128, context 48 steps
outcome model MLP 256, horizon 48 steps, targets dT / 5 degC and energy / 16.7 Wh/m2, lambda_E = lambda_C = 1
flow generator MLP 4 x 512, held plan x = [w / 30, mode] in [-1, 1]^6, HVAC-off source + 0.25 N(0, I)

Metrics: comfort penalty = mean over steps and conditioned zones of (T - 21)^2 [degC^2]; energy = HVAC electricity + gas [kWh]; normalized return = return / tuned-RBC return on the same building and task (1.0 = RBC, lower is better).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support