Instructions to use rodneyslafuente/decision-model-rl-overcooked with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use rodneyslafuente/decision-model-rl-overcooked with PEFT:
from peft import PeftModel from transformers import AutoModelForSequenceClassification base_model = AutoModelForSequenceClassification.from_pretrained("rodneyslafuente/openjev-v5-0.8b") model = PeftModel.from_pretrained(base_model, "rodneyslafuente/decision-model-rl-overcooked") - Notebooks
- Google Colab
- Kaggle
A decision model that learned to cook and cooperate
Get the code and instructions, or read the write-up and watch the video.
Two independently acting chefs share this policy and serve six soups in 512 ticks in PufferLib's Cramped Room. Each sees its public observation expressed in text, game rules, and eight recent action outcomes. The policy scores stay, up, down, left, right, and interact. Each selection executes one native action, without a planner, pathfinder, action macro, or assigned role.
This repository contains LoRA adapters and the trained original NLI classifier head, not a standalone base model. The root is cooperative checkpoint 220. single-chef/ contains checkpoint 330, which initialized cooperative training. Both use OpenJev v5's 0.8B checkpoint. The ancestry is Qwen3.5-0.8B, then OpenJev v5 0.8B, then this cooking adapter. The adapter points to an unchanged, attributed mirror of the exact upstream subfolder. A separate base repository avoids inheriting the unrelated 4B ancestry from OpenJev's shared repository.
Use
The code repository downloads the exact base model, builds the pinned native environment, evaluates the policy, and renders video with the original prompts and scores.
python download.py
python evaluate.py
python render.py runs/evaluation.json 220 --folder videos
For standard Transformers and PEFT loading:
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from peft import PeftModel
base_id = "rodneyslafuente/openjev-v5-0.8b"
revision = "61af02b317b1e462a7594721823c234974ec1a5e"
tokenizer = AutoTokenizer.from_pretrained(base_id, revision=revision)
base = AutoModelForSequenceClassification.from_pretrained(
base_id, revision=revision, dtype=torch.bfloat16
).to("cuda")
base.config.get_text_config().pad_token_id = tokenizer.pad_token_id
model = PeftModel.from_pretrained(base, "rodneyslafuente/decision-model-rl-overcooked").eval()
The adapter metadata also supports automatic PEFT base resolution through the pinned 0.8B mirror. Use the repository's pinned dependencies. policy.yaml contains the prompt rules and training settings. The underlying NLI template comes from the base model config. For each proposed button, score the hypothesis Your next action should be {action}., take the entailment probability (class 1), and normalize across the six candidates. These are action-selection scores, not calibrated probabilities of success. The code uses optimized shared-prefix scoring for the reported evaluation.
Training and result
I trained text-backbone LoRA adapters with rank 16 and alpha 32 and fully trained the existing three-class NLI head. The adapter contains 10,825,728 trained parameters. No action-specific head was added. The algorithm is a clipped group-relative policy gradient with discounted reward-to-go, adaptive entropy, observation novelty, and a repeated-no-op penalty. The game provides rewards, including shaping for adding ingredients, starting cooking, and plating. There are no demonstrations or teacher model.
Training used one 16 GB RTX 5080. The successful single-chef stage took about 19.5 hours, followed by about 32.6 hours of cooperative training to checkpoint 220, excluding an earlier unsuccessful run. Cooperative updates collect 8,192 agent decisions across eight kitchens. Learning rates are 5e-6 for LoRA and 1e-5 for the head. The recorded checkpoint-220 update peaked at 7.45 GiB allocated VRAM.
The six-soup result is one greedy evaluation, seed 0, selected as the best checkpoint using that same seed. It is not a multi-seed average or held-out test. The two chefs share weights, but each chooses independently from its own observation. Cooperation appears in the trace as one chef plating or serving while the other adds onions. A trial on another layout did not transfer successfully. Retention on unrelated NLI tasks and visual control were not evaluated for these trained weights.
evaluation.json records all 512 ticks, both chefs' exact states, action scores, and outcomes. provenance.json records the base/environment revisions and the original adapter checksum. Training changed settings during development. The final configs support evaluation and continued training, but do not capture every historical intervention.
Attribution
Base: OpenJev, by AlexWortega, with a Qwen3.5-0.8B backbone. Environment: PufferLib. This release distributes the trained adapter under MIT. The base weights and dependencies retain their upstream licenses. The write-up includes the full method and references.
- Downloads last month
- -
Model tree for rodneyslafuente/decision-model-rl-overcooked
Base model
Qwen/Qwen3.5-0.8B-Base