You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Contains weights derived from an unreleased base (Molmo2-4B-SFT). Research use.

Log in or Sign Up to review the conditions and access this model content.

molmo2-codec Stage-2 V4 (trained codec video model)

Stage-2 SFT checkpoints for the AdaCodec-on-Molmo2 video pipeline, V4 (trainable E_P P-frame ViT branch). Video is ingested as a codec representation β€” I-frames: 81 tokens, P-frames: 16 tokens via the trained E_P β€” at ~33% of dense's visual-token budget (3x compression).

Results β€” MLVU multiple-choice (n=100, official dense-vs-codec eval, same 100 examples)

checkpoint codec acc token budget
v2b (frozen P-tokenizer, 0.49M) 35.0% ~25%
stage2_v4_step500.pt 46.0% 33.2%
stage2_v4_step1200.pt 52.0% 33.2%
dense (per-frame RGB, reference) 71.0% 100%

The trainable ~381M E_P branch (vs the old frozen 0.49M P-tokenizer) is the key change: codec accuracy went 35% -> 46% (500 steps) -> 52% (1200 steps) and was still climbing.

Files

  • stage2_v4_step500.pt, stage2_v4_step1200.pt β€” consolidated {model, proj} state dicts (~19GB fp32, full 4.85B model incl. the trained E_P). Load with codec/codec_eval.py --ckpt <file> (E_P auto-attaches because the ckpt carries codec_ptok.vit_p.*).

Code: https://github.com/weikaih04/molmo2-codec (branch adacodec). Stage-1 E_P: https://huggingface.co/weikaih/molmo2-codec-stage1-v4

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support