Instructions to use sinnybb/exact_random_selfdistill_grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sinnybb/exact_random_selfdistill_grpo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="sinnybb/exact_random_selfdistill_grpo") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("sinnybb/exact_random_selfdistill_grpo") model = AutoModelForMultimodalLM.from_pretrained("sinnybb/exact_random_selfdistill_grpo", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sinnybb/exact_random_selfdistill_grpo with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sinnybb/exact_random_selfdistill_grpo" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sinnybb/exact_random_selfdistill_grpo", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/sinnybb/exact_random_selfdistill_grpo
- SGLang
How to use sinnybb/exact_random_selfdistill_grpo with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sinnybb/exact_random_selfdistill_grpo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sinnybb/exact_random_selfdistill_grpo", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sinnybb/exact_random_selfdistill_grpo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sinnybb/exact_random_selfdistill_grpo", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use sinnybb/exact_random_selfdistill_grpo with Docker Model Runner:
docker model run hf.co/sinnybb/exact_random_selfdistill_grpo
Teacher exact39K random curriculum Stage3 โ GRPO
This repository contains the final, merged GRPO checkpoint of
our_stage1_ep2_teacher_exact39k_random42_stage3_exact39k_ep1.
The starting checkpoint is the Teacher exact39K Stage3 model: Stage1 training for 2 epochs, a seed-42 random curriculum of three 13K subsets with 2 epochs per subset, followed by Stage3 training on teacher-correct examples for 1 epoch.
The published checkpoint is GRPO global step 12, the final saved step of the completed run. The four distributed actor shards were merged into four Hugging Face safetensors shards. Tokenizer and image processor files are included.
RL configuration
| Setting | Value |
|---|---|
| Algorithm | GRPO |
| Training data | Thyme-RL, maximum 3,200 training examples |
| Epochs | 1 |
| GPUs | 4 |
| Latent size | 10 |
| Rollouts per prompt | 8 |
| Sampling temperature | 0.5 |
| Learning rate | 1e-6 |
| KL coefficient | 0.01 |
| Monet RL sigma | 10.0 |
| Final saved step | 12 |
The architecture is Qwen2.5-VL. Monet latent reasoning requires the Monet inference/runtime implementation used by the training project.
- Downloads last month
- -