GR00T-N1.7-LIBERO-w4a4

nvidia/GR00T-N1.7-LIBERO quantized with FoldQuantVLA: 4-bit weights and 4-bit activations in every LLM and action-head projection. Everything else (vision encoder, embeddings, norms, small encoders) is unchanged.

The repo has one folder per LIBERO suite (libero_spatial, libero_object, libero_goal, libero_10), like the base repo. Each folder is a full checkpoint for that suite.

Base model nvidia/GR00T-N1.7-LIBERO
Scheme --llm-scheme w4a4_srg, action head w4a4_shg
Calibration 128 frames from LIBERO demonstrations, seed 0
Size about 4 GB per suite

Quantized layers store integer weights (qweight, two 4-bit values per byte, or int8) and one fp32 scale per output row (weight_scale). foldquant_quant.json lists the quantized layers and their settings.

1. Install

git clone https://github.com/cair-vinuni/FoldQuantVLA.git
cd FoldQuantVLA
git submodule update --init third_party/cutlass models/groot_n1_7/external_dependencies/LIBERO
cd models/groot_n1_7
uv sync
source .venv/bin/activate
uv pip install -e ../..
uv pip install "robosuite==1.4.0" "mujoco==2.3.7" bddl easydict hydra-core einops termcolor thop gym
hf auth login   # the backbone needs access to the gated nvidia/Cosmos-Reason2-2B

The commands below run from FoldQuantVLA/models/groot_n1_7 with this environment active.

2. Download

hf download vrfai/GR00T-N1.7-LIBERO-w4a4 --local-dir checkpoints/GR00T-N1.7-LIBERO-w4a4
# or a single suite:
hf download vrfai/GR00T-N1.7-LIBERO-w4a4 --include "libero_10/*" --local-dir checkpoints/GR00T-N1.7-LIBERO-w4a4

3. Run in PyTorch

Evaluate on LIBERO (10 tasks x 20 episodes, on the matching suite, the paper's P3 protocol):

MUJOCO_GL=egl python -m foldquant_integration.eval_libero --protocol p3 \
    --model-path checkpoints/GR00T-N1.7-LIBERO-w4a4/libero_10 --suites libero_10 --max-episode-steps 720 --output results/w4a4/libero_10

Results go to results/w4a4/libero_10/summary.json. For another suite, change both the folder and --suites (e.g. libero_goal). Add --n-episodes 1 for a quick check.

Start a policy server (same interface as the upstream GR00T server):

python -m foldquant_integration.serve --model-path checkpoints/GR00T-N1.7-LIBERO-w4a4/libero_10 --embodiment-tag libero_panda --port 5555

In PyTorch the quantized layers are computed from the stored integers in fp32. The outputs match the TensorRT engines, but this path is not faster than bf16. Use TensorRT for speed.

4. Run with TensorRT

Build the INT4/INT8 kernels and the engines on the machine that will run them (engines depend on the GPU and the TensorRT version):

python -m foldquant.kernels build
python -m foldquant_integration.export --model-path checkpoints/GR00T-N1.7-LIBERO-w4a4/libero_10 --embodiment-tag libero_panda --output-dir exports/w4a4/onnx
python -m foldquant_integration.build_engines --onnx-dir exports/w4a4/onnx --engine-dir exports/w4a4/engines
MUJOCO_GL=egl python -m foldquant_integration.eval_libero --protocol p3 \
    --model-path checkpoints/GR00T-N1.7-LIBERO-w4a4/libero_10 --engine-dir exports/w4a4/engines --suites libero_10 --max-episode-steps 720 --output results/w4a4/libero_10_trt

4-bit matrix multiplies run natively on sm_80 / sm_86 / sm_87 (Jetson AGX Orin) / sm_89 GPUs. H100 (sm_90) has no 4-bit tensor-core path and runs them on its 8-bit units.

Results

Measured on one H100. Action cosine: decoded action chunk vs the bf16 model on held-out LIBERO frames (1.0 = identical). Success rate: LIBERO, P3 protocol (10 tasks x 20 episodes, on the matching suite, seed 7, 720 steps).

suite action cosine success rate (w4a4) success rate (bf16 base)
libero_spatial 0.9983 195/200 (97.5%) 196/200 (98.0%)
libero_object 0.9949 196/200 (98.0%) 198/200 (99.0%)
libero_goal 0.9958 186/200 (93.0%) 192/200 (96.0%)
libero_10 0.9987 175/200 (87.5%) 178/200 (89.0%)
all 752/800 (94.0%) 764/800 (95.5%)

Latency and memory on one H100, batch 1, TensorRT engines (this checkpoint vs the bf16 base built the same way). H100 has no 4-bit tensor cores, so the 4-bit layers run through a slower path there. For latency on GPUs with 4-bit tensor cores (RTX 40 series, Jetson AGX Orin), see the FoldQuantVLA paper.

bf16 engines w4a4 engines
query latency (ms) 26.0 114.0
engine files (MiB) 6197 3520
GPU memory while serving (MiB) 7559 11337
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for vrfai/GR00T-N1.7-LIBERO-w4a4

Finetuned
(5)
this model

Collection including vrfai/GR00T-N1.7-LIBERO-w4a4