CS2 Match-Level Hierarchical World Model (scale300)

A causal, action-conditioned latent world model over frozen V-JEPA2 embeddings of single-POV Counter-Strike 2 gameplay. Sequences span a player's whole match (rounds concatenated, round-number embedding per token), so the belief state can carry cross-round context. Trained on 300 matches of RekaAI/CS2-10k (mirage + dust2), 6.94M one-second windows at 0.5s stride.

Screen-only contract: the encoder sees POV pixels with no HUD, radar, or parsed game state. Positions and actions are used as training targets and for evaluation, never as inference inputs.

Files

file what it is
hwm_ckpt_scale300.pt the world model (12.25M params) + latent norm stats + action norm
demo_heads.pt readout heads on the same latents: position, position@+8s, behaviour cloning (9 keys), 30s zone intent, zone-transition prior
reports/*.json every number below, as produced by the scripts

Results (match-disjoint holdout, 450 sequences)

Latent MSE against the true future latent, versus a persistence ("nothing changes") baseline:

horizon model persistence
0.5 s 0.279 0.417
1 s 0.436 0.765
2 s 0.617 1.091
4 s 0.765 1.369
8 s 0.882 1.608
30 s 0.976 1.893

The limitation that matters most

Shuffling the action input and re-measuring gives the share of the model's headroom over persistence that actually comes from knowing the action:

horizon 0.5 s 1 s 2 s 4 s 8 s 30 s
action share of headroom 24.4% 11.6% 6.6% 3.1% 1.2% βˆ’0.0%

At 30 s the model is provably action-blind, and its MSE of 0.976 on ~unit-variance latents means it is close to a mean predictor there. It beats persistence at 30 s only because persistence is a poor baseline at that range.

Structurally: forward(z, a, rnd) is causally masked, so token t sees actions only up to t and each horizon head predicts zt+h directly β€” there is no slot for a candidate future action sequence. Counterfactual "what if I do X" requires autoregressive rollout through the h=1 head and is only meaningful for ~0.5–2 s. Macro intent must come from the destination/intent heads, not rollouts.

Reproduce with scripts/hwm/hwm_action_sensitivity.py.

30 s intent probe

features top-1 top-3 moved top-1
latent 0.420 0.781 0.269
HWM context 0.413 0.776 0.303
latent + context 0.407 0.771 0.299
persistence baseline 0.371 β€” β€”
zone-transition prior (moved subset) β€” β€” 0.318

Latent top-1 clears the persistence baseline. On the moved subset (63% of cases, where the player actually changes zone β€” the ones that matter for next-action) all three feature sets still lose to a trivial zone-transition prior. beats_transition_prior is false everywhere. Scaling 100 β†’ 300 matches narrowed that gap but did not close it, and improved the dynamics model only ~2%, so scaling in this range is sublinear.

Demo: verified on holdout video

scripts/hwm/demo_pov_video.py replays a holdout match as real video through the frozen encoder and this checkpoint, with the action input supplied by the BC head reading the same pixels. Ground truth is overlaid only to check.

metric result baseline
self-localization (median) 1.99 m β€”
8 s position forecast 6.03 m 9.68 m stand-still
30 s intent top-3 60.3% 55.8% zone-transition prior
2 s counterfactual, true action closest 46.2% 16.7% chance

Caveats, in the interest of not overstating: the 8 s figure shown is a direct forecast head β€” decoding the HWM's own h=16 latent scores slightly worse (6.53 m), so the world model is not what drives that panel. The counterfactual is scored in latent space; scored on map positions it falls to 18.6%, i.e. chance, because the position head's ~99 u error swamps the 238 u candidate spread. It is an aggregate claim, not a per-frame one. The position head is unconstrained and emits off-map coordinates on ~1% of frames.

Intended use

Research on latent world models and intent prediction from egocentric video. Not a cheat, not an aim assist: it consumes recorded POV video offline and predicts where a player will be, at metre-to-zone granularity.

Provenance

Derived from RekaAI/CS2-10k; that dataset's terms apply upstream. Encoder is facebook/vjepa2-vitl-fpc64-256, frozen. Match split is deterministic (md5 of match_id), so the holdout is reproducible. The checkpoint itself is not bit-reproducible: training seeds the RNGs but GPU nondeterminism means a re-run yields a statistically equivalent, not identical, model.

Code: https://github.com/cristibctr/HWM-CS2

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support