You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Llama-3.2-3B-Instruct-SFT-ScienceWorld

PRIVATE — internal research checkpoint. Do not redistribute.

Task ScienceWorld
Stage SFT (behavior-cloning) initialization
Base meta-llama/Llama-3.2-3B-Instruct (Built with Llama)
Init for this run models/llama-3.2-3b-instruct
Source checkpoint rl/ckpts/sw_sft_llama3.2-3b/global_step_198 (global_step 198)
Format bf16 safetensors, merged from hf-dir -> bf16 cast
Params 3,212,749,824
Weight drift vs. init (mean rel-L2) 4.69e-03 (max 9.40e-03, 226 / 255 tensors changed)
Reload bit-exact check True

Optimizer state is not included (inference/eval only; the fp32 FSDP shards stay on HiPerGator /blue). Load with AutoModelForCausalLM.from_pretrained(..., torch_dtype=torch.bfloat16) or vLLM; the chat template is the stock Llama-3 template shipped in chat_template.jinja.

Licensed under the Llama 3.2 Community License; derivative of Meta Llama.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wayne377/Llama-3.2-3B-Instruct-SFT-ScienceWorld

Finetuned
(2016)
this model