Latent Student Spatial Forcing

This is the standalone Stage 4 inference package for the Latent Student. The Stage 4 LoRA adapter has been merged into shreethar/LatentStudent-ckpt-400.

Included components

  • Merged Qwen3.5 vision-language model weights
  • Processor and tokenizer
  • spatial_parameters.pt: five learned spatial-slot embeddings and the Stage 4 waypoint MLP
  • latent_student_config.json: packaging and provenance metadata
  • stage4_config.json: training configuration, when present in the checkpoint

VGGT and the Spatial Forcing projection head were training-only supervision components. They are not needed for waypoint inference.

Provenance

  • Stage 2 model: shreethar/LatentStudent-ckpt-400
  • Stage 4 checkpoint: stage4_partial_run_2/step_002650
  • Checkpoint selection: best_checkpoint.json
  • Checkpoint step: 002650
  • Loss weights: alpha=1.0, beta=3.0, gamma=0.025

Loading for waypoint inference

Use the project's LatentStudent wrapper so the spatial slots and waypoint head are restored alongside the merged VLM:

from transformers import AutoTokenizer
from train.stage4.checkpointing import load_latent_student_checkpoint

repo_id = "shreethar/Latent-Student-Spatial-Forcing"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
end_think_token_id = tokenizer.convert_tokens_to_ids("</think>")

student = load_latent_student_checkpoint(
    checkpoint=repo_id,
    end_think_token_id=end_think_token_id,
    trainable=False,
    M=6,
    K=5,
)
student.eval()

Loading only with AutoModelForImageTextToText restores the merged VLM but not the external spatial slots or waypoint head. Use the wrapper above for the complete Latent Student behavior.

Downloads last month
2
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shreethar/Latent-Student-Spatial-Forcing

Finetuned
(1)
this model