GON: Generalizable Outcome Networks
Code for GON: Generalizable Outcome Networks for Learning Action Effects in Physical Systems (anonymous submission).
GON learns how HVAC action sequences change a building's future from matched counterfactual branches: simulations that start from the same building state under the same weather and occupancy but apply different plans. An outcome model trained with a matched-branch loss predicts temperature and energy for a proposed plan; a conditional flow-matching generator proposes many plans for a requested temperature change; at deployment the outcome model scores the candidates and the first action of the best plan is executed, every 5 minutes. One frozen controller is evaluated zero-shot on 100 unseen Building2Building buildings.
Layout
gon/
constants.py observation/action layout, normalization, delivered-heat actuation map
env.py Building2Building env with the 5-minute step fix, episode loop, metrics
baselines.py tuned RBC (normalization anchor, logging policy) and PI controller
branching.py logged episodes, plan family, matched counterfactual branching (replay + diverge)
data.py history windows, matched-pair dataset, flow-generator items
models.py history encoder, outcome model + matched-branch loss, flow generator, checkpoints
controller.py the GON deployment controller
utils.py paths, splits, process pool
scripts/ collect_logs -> collect_branches -> train_encoder -> train_outcome -> train_flow -> evaluate
Setup
pip install -r requirements.txt
pip install -e <path-to-building2building> # EnergyPlus-based benchmark
export PYTHONPATH=$PWD:<path-to-building2building>:$PYTHONPATH # `baselines` (tuned RBC) lives in that repository
Main settings
| component | setting |
|---|---|
| environment | Building2Building OfficeSmall, winter (Jan 1 - Mar 31), 5-min step (minimum_system_timestep = 0), 25 920 steps |
| tasks | task_const_e0 (comfort-focused, w_E = 0), task_const_e05 (energy-focused, w_E = 0.5), setpoint 21 degC |
| encoder | causal transformer, 4 layers, 4 heads, width 128, context 48 steps |
| outcome model | MLP 256, horizon 48 steps, targets dT / 5 degC and energy / 16.7 Wh/m2, lambda_E = lambda_C = 1 |
| flow generator | MLP 4 x 512, held plan x = [w / 30, mode] in [-1, 1]^6, HVAC-off source + 0.25 N(0, I) |
Metrics: comfort penalty = mean over steps and conditioned zones of (T - 21)^2 [degC^2]; energy = HVAC electricity + gas [kWh]; normalized return = return / tuned-RBC return on the same building and task (1.0 = RBC, lower is better).