Holographic Entropy Cone β€” RL policies and inverse map

Model weights from the entropy-cone research programme: reinforcement-learning policies that realise target entropy vectors as weighted graphs, and a supervised inverse map that predicts a graph's edge support directly from a target entropy vector.

Code: https://github.com/Jaeha0526/entropyconeRL

Related paper (earlier RL stage): Discovering the Holographic Entropy Cone via Reinforcement Learning β€” Jaeha Lee, Temple He, Hirosi Ooguri (JHEP 2026).

Why the inverse map matters

For the n=6 cone, continuous RL policies plateau around cosine 0.995 to a target ray β€” close, but never exact, because the last mile is a combinatorial problem (which ~16 of 190 edges are present), not a precision problem.

The inverse map attacks exactly that: an MLP predicts the support from the 63-dimensional entropy vector, then the weights follow from a constrained solve on that support, and the result is verified exactly through the cut matrix.

Using this, 1,229 of 4,161 catalogued extreme rays (29.5%) were given exact, machine-verified graph realisations, including two candidate rays not present in the reference catalogue. Held-out performance matched training performance (31.2% vs 27.0%), indicating the map learned structure rather than memorising.

Contents

inverse_map/
  inverse_map_v2_final.pt              MLP 63 -> support logits (190) + log-weights
                                       the model behind the 1,229 exact certificates

n6_policy/
  n6_gnn_v2tree_N20_seed42_iter8500_best.pt   best n=6 policy (mean gap 0.004309)
  n6_gnn_v2tree_N20_seed42_final.pt           same run, final iterate
  n6_raytgt_ft_seed42_final.pt                ray-targeted fine-tune endpoint
  n6_gnn_v1_N18_seed42_final.pt               N=18 variant
  n6_gnn_v1_h192_seed42_final.pt              wider hidden (192)
  n6_gnn_v1_seed42_h200_final.pt              wider hidden (200)
  n6_gnn_v2tree_seed42_final.pt               tree-sampling variant

n5_policy/
  n5_gnn_v1_seed{42,123,456}_final.pt         n=5 GNN policies, three seeds
  n5_gnn_v2_seed42_final.pt                   v2 recipe
  n5_mlp_aug_v1_seed42_final.pt               MLP baseline with augmentation

Caveat on the ray-targeted fine-tune

n6_raytgt_ft_seed42_final.pt is the endpoint of a negative experiment: training with 30% of targets drawn from catalogue extreme rays moved the median gap only 4.8e-3 β†’ 4.3e-3 and produced no exact realisations. It is included for reproducibility, not because it improves on the baseline.

Loading

from huggingface_hub import hf_hub_download
import torch

path = hf_hub_download("JaehaL/entropy-cone-models",
                       "inverse_map/inverse_map_v2_final.pt")
ckpt = torch.load(path, map_location="cpu")   # dict with 'net', 'arch', ...

The n=6 policies expect the HECVecEnv observation layout (n=6, N=20, s_dim 63, w_dim 190) defined in the repository.

Provenance

Trained on the Caltech Resnick HPC, 2025–2026. Training configs, the exact finisher, verification scripts, and experiment notes are in the GitHub repository; this archive holds only the weights.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading