Holographic Entropy Cone β RL policies and inverse map
Model weights from the entropy-cone research programme: reinforcement-learning policies that realise target entropy vectors as weighted graphs, and a supervised inverse map that predicts a graph's edge support directly from a target entropy vector.
Code: https://github.com/Jaeha0526/entropyconeRL
Related paper (earlier RL stage): Discovering the Holographic Entropy Cone via Reinforcement Learning β Jaeha Lee, Temple He, Hirosi Ooguri (JHEP 2026).
Why the inverse map matters
For the n=6 cone, continuous RL policies plateau around cosine 0.995 to a target ray β close, but never exact, because the last mile is a combinatorial problem (which ~16 of 190 edges are present), not a precision problem.
The inverse map attacks exactly that: an MLP predicts the support from the 63-dimensional entropy vector, then the weights follow from a constrained solve on that support, and the result is verified exactly through the cut matrix.
Using this, 1,229 of 4,161 catalogued extreme rays (29.5%) were given exact, machine-verified graph realisations, including two candidate rays not present in the reference catalogue. Held-out performance matched training performance (31.2% vs 27.0%), indicating the map learned structure rather than memorising.
Contents
inverse_map/
inverse_map_v2_final.pt MLP 63 -> support logits (190) + log-weights
the model behind the 1,229 exact certificates
n6_policy/
n6_gnn_v2tree_N20_seed42_iter8500_best.pt best n=6 policy (mean gap 0.004309)
n6_gnn_v2tree_N20_seed42_final.pt same run, final iterate
n6_raytgt_ft_seed42_final.pt ray-targeted fine-tune endpoint
n6_gnn_v1_N18_seed42_final.pt N=18 variant
n6_gnn_v1_h192_seed42_final.pt wider hidden (192)
n6_gnn_v1_seed42_h200_final.pt wider hidden (200)
n6_gnn_v2tree_seed42_final.pt tree-sampling variant
n5_policy/
n5_gnn_v1_seed{42,123,456}_final.pt n=5 GNN policies, three seeds
n5_gnn_v2_seed42_final.pt v2 recipe
n5_mlp_aug_v1_seed42_final.pt MLP baseline with augmentation
Caveat on the ray-targeted fine-tune
n6_raytgt_ft_seed42_final.pt is the endpoint of a negative experiment:
training with 30% of targets drawn from catalogue extreme rays moved the median
gap only 4.8e-3 β 4.3e-3 and produced no exact realisations. It is included for
reproducibility, not because it improves on the baseline.
Loading
from huggingface_hub import hf_hub_download
import torch
path = hf_hub_download("JaehaL/entropy-cone-models",
"inverse_map/inverse_map_v2_final.pt")
ckpt = torch.load(path, map_location="cpu") # dict with 'net', 'arch', ...
The n=6 policies expect the HECVecEnv observation layout (n=6, N=20,
s_dim 63, w_dim 190) defined in the repository.
Provenance
Trained on the Caltech Resnick HPC, 2025β2026. Training configs, the exact finisher, verification scripts, and experiment notes are in the GitHub repository; this archive holds only the weights.