Instructions to use rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-E2B-it") model = PeftModel.from_pretrained(base_model, "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft") - Transformers
How to use rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft
- SGLang
How to use rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft with Docker Model Runner:
docker model run hf.co/rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft
Procurement Function Calling — Gemma 4 E2B LoRA SFT
This repository contains a LoRA adapter for google/gemma-4-E2B-it, supervised
fine-tuned for multilingual, multi-turn procurement tool calling. The model
searches suppliers, requests quotes, evaluates delivery options and submits a
procurement plan while respecting budget, delivery, reliability and compliance
constraints.
Results
Every run evaluates all 400 scenario–prompt pairs in the held-out test split, including unseen templates 16–25. All configurations use the same seed and test examples. Values are the mean ± sample standard deviation across three repeated inference runs.
| Model | Thinking | Runs | Task success | Average return | Mean episode latency |
|---|---|---|---|---|---|
| Base | Disabled | 3 | 16.58% ± 0.29 pp | 0.230 ± 0.010 | 15.75 ± 0.14 s |
| Base | Enabled | 3 | 40.33% ± 1.66 pp | 0.476 ± 0.017 | 95.36 ± 2.04 s |
| SFT | Disabled | 3 | 42.50% ± 0.35 pp | 0.567 ± 0.005 | 20.31 ± 0.19 s |
Against the non-thinking base, SFT improves mean task success by 25.92 percentage points and mean return by 0.338. It also slightly exceeds the thinking-enabled base while using about one fifth of its mean episode latency.
The 48-step job processed 1,472,140 non-padding input tokens across 384 examples. Of these, 188,297 assistant tool-call tokens carried loss. Training took 65 minutes on one NVIDIA A40 GPU. The run was deliberately compute-bounded and was not optimized for maximum environment return or task success.
Loading
from peft import AutoPeftModelForCausalLM
from transformers import AutoTokenizer
model_id = "rubenbalbastre/procurement-function-calling-gemma-4-E2B-it-sft"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoPeftModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto",
)
This repository contains the adapter rather than the complete base-model weights. Use the tokenizer's chat template and provide the procurement tool schemas when constructing requests.
Training
- Base model:
google/gemma-4-E2B-it - Dataset:
rubenbalbastre/supply-chain-tool-calling - Method: assistant-only LoRA supervised fine-tuning with PEFT and TRL
- LoRA rank: 16
- LoRA alpha: 32
- LoRA dropout: 0.05
- Optimizer steps: 48
- Thinking during SFT: disabled
Limitations
- Results are specific to the included synthetic procurement environment, tools and verifier.
- Preferred-supplier fallback remains the weakest task family.
- The model can emit invalid calls, select infeasible plans or fail to complete longer trajectories.
- Do not use its output as an autonomous purchasing decision without external validation.
Source
Training, data-generation and evaluation code is available in the
sft-tool-calling
repository.
- Downloads last month
- 29