Instructions to use annajuliaasf/gemma-4-e2b-tool-use-ptbr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use annajuliaasf/gemma-4-e2b-tool-use-ptbr with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/gemma-4-e2b-it-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "annajuliaasf/gemma-4-e2b-tool-use-ptbr") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use annajuliaasf/gemma-4-e2b-tool-use-ptbr with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for annajuliaasf/gemma-4-e2b-tool-use-ptbr to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for annajuliaasf/gemma-4-e2b-tool-use-ptbr to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for annajuliaasf/gemma-4-e2b-tool-use-ptbr to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="annajuliaasf/gemma-4-e2b-tool-use-ptbr", max_seq_length=2048, )
Gemma 4 E2B: LoRA adapter for tool calling in Portuguese
A LoRA adapter that teaches unsloth/gemma-4-E2B-it to follow a function calling
protocol in Portuguese: deciding whether a question requires a tool, choosing which one,
filling in the arguments in JSON, and answering in prose when no tool is needed.
Study project. The training data is synthetic and the 30 tools are fictional; none of them corresponds to a real API.
What it produces
Every assistant turn comes out delimited, which lets a program decide without ambiguity whether to execute a tool or show the text:
<tool_call>
{"nome_tool": "create_budget", "argumentos": {"income": 5000, "expenses": [...]}}
</tool_call>
<final_answer>
Melhorar as habilidades de resolução de conflitos envolve desenvolver comunicação eficaz...
</final_answer>
The tool's return is handed back to the model in <tool_result>...</tool_result>, and the
model stops generating after emitting the call. That behaviour is what gives a real
agent its turn to execute the API.
How to use
Load the model with the snippet from the Use this model button at the top of this page.
Training used Unsloth, but inference does not need it: this is a plain PEFT adapter, so
peft and bitsandbytes are enough. Verified on a free Colab T4.
Once you have model, this is the part the button cannot generate, because it depends on how
the model was trained:
import torch
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("annajuliaasf/gemma-4-e2b-tool-use-ptbr")
dev = next(model.parameters()).device
INSTRUCOES = """<instrucoes>
Você é um assistente de IA com acesso a um conjunto de ferramentas.
Seu objetivo é responder às perguntas do usuário de forma correta e útil.
Como proceder:
1. Analise o pedido do usuário e entenda a intenção.
2. Decida se precisa de ferramenta:
- Se você consegue responder com o que já sabe, responda direto.
- Se precisa de informação externa ou de executar uma ação, use uma ferramenta.
3. Se for usar uma ferramenta, escolha a mais adequada entre as disponíveis
e extraia os argumentos a partir do pedido do usuário.
4. Para chamar uma ferramenta, escreva a chamada dentro de <tool_call>.
O conteúdo deve ser um objeto JSON:
<tool_call>
{"nome_tool": "nome_exato_da_ferramenta", "argumentos": {"parametro": "valor"}}
</tool_call>
5. O resultado da ferramenta será devolvido a você dentro de <tool_result>.
Interprete esse resultado para responder ao pedido original.
6. Sua resposta final ao usuário deve estar sempre dentro de <final_answer>,
em linguagem natural. Se a ferramenta retornou erro, explique isso ao usuário.
<final_answer>
Sua resposta ao usuário.
</final_answer>
</instrucoes>"""
FERRAMENTAS = """<ferramentas>
nome:get_traffic_info, descrição: Retorna informações de trânsito para uma rota específica., parâmetros: [{'nome': 'origin', 'tipo': 'string', 'obrigatorio': True}, {'nome': 'destination', 'tipo': 'string', 'obrigatorio': True}]
nome:get_stock_price, descrição: Retorna o preço atual de uma ação específica., parâmetros: [{'nome': 'symbol', 'tipo': 'string', 'obrigatorio': True}]
nome:book_hotel, descrição: Reserva um quarto de hotel com as opções especificadas., parâmetros: [{'nome': 'hotel_name', 'tipo': 'string', 'obrigatorio': True}, {'nome': 'check_in', 'tipo': 'string', 'obrigatorio': True}, {'nome': 'check_out', 'tipo': 'string', 'obrigatorio': True}]
nome:create_budget, descrição: Cria um orçamento com base nas receitas e despesas., parâmetros: [{'nome': 'income', 'tipo': 'number', 'obrigatorio': True}, {'nome': 'expenses', 'tipo': 'array', 'obrigatorio': True}]
</ferramentas>"""
mensagens = [
{"role": "system", "content": INSTRUCOES + "\n\n" + FERRAMENTAS},
{"role": "user", "content": "Quero saber o valor atual da ação da Microsoft."},
]
texto = tokenizer.apply_chat_template(mensagens, tokenize=False, add_generation_prompt=True)
entrada = tokenizer(texto, return_tensors="pt", add_special_tokens=False).to(dev)
with torch.no_grad():
saida = model.generate(**entrada, max_new_tokens=800, do_sample=False)
print(tokenizer.decode(saida[0][entrada["input_ids"].shape[1]:], skip_special_tokens=True))
Expected output:
<tool_call>
{"nome_tool": "get_stock_price", "argumentos": {"symbol": "MSFT"}}
</tool_call>
Four details that matter:
- Both blocks are required, and this was measured. With the same catalog but no
<instrucoes>block, the model still picks the right tool and the right argument, but it emits<get_stock_price symbol="MSFT"/>, Gemma's own native format, instead of the<tool_call>protocol. The fine-tuning taught it to follow the instruction very well, which is not the same as no longer needing it. Keep the instruction block verbatim. - The catalog is the interface. The model picks from the tools listed in the system prompt, so it can work with tools it never saw in training. That is what conditions B and D below measure. No catalog means nothing to pick from.
- Swap the tools for your own. The four above are a sample of the 30 fictional tools used
in training. List whichever ones your application actually has, one per line, in that exact
nome:..., descrição: ..., parâmetros: [...]shape, wrapped in<ferramentas>tags. add_special_tokens=False:apply_chat_templatealready inserts the<bos>. Letting the tokenizer insert another one degrades the result without raising an error.do_sample=False: a tool call needs exact JSON, not creativity.
Training
| base | unsloth/gemma-4-E2B-it (5.12B) |
| base actually loaded | unsloth/gemma-4-e2b-it-unsloth-bnb-4bit, the pre-quantized 4-bit copy |
| method | QLoRA · r=8, lora_alpha=8, lora_dropout=0 |
| layers | attention + MLP; finetune_vision_layers=False |
| data | 386 training examples, 47 validation |
| batch | 1 × 8 gradient accumulation (effective 8) |
| epochs | 3 (147 steps) |
| lr | 2e-4, linear scheduler |
| optimizer | adamw_8bit |
| loss | train_on_responses_only, assistant turns only |
| hardware | 1× Tesla T4 (free Colab), 34min21, float32 |
load_in_4bit=True makes Unsloth silently swap the base for its pre-quantized 4-bit copy,
so adapter_config.json records unsloth/gemma-4-e2b-it-unsloth-bnb-4bit rather than the
16-bit repository named above. That is the base the snippet under "How to use" loads, and it
requires bitsandbytes and a GPU.
Final loss: train 0.053 / validation 0.488. The validation loss never went back up; it flattened.
| epoch | validation at the end | gain |
|---|---|---|
| start (step 10) | 0.939 | n/a |
| 1 (step ~50) | 0.528 | −0.41 |
| 2 (step ~100) | 0.492 | −0.036 |
| 3 (step 147) | 0.488 | −0.004 |
Each epoch returned ten times less than the previous one, so 1 to 2 epochs would have been enough.
Evaluation
51 test traces never seen during training, under four conditions. The base model received the same system prompt, with the full instructions: the comparison is instruction against training, not "knows the format" against "never saw it".
| condition | decided right | emitted <tool_call> |
valid JSON | right tool | params ok | types ok |
|---|---|---|---|---|---|---|
| A original | 51/51 = 100% | 37/37 | 37/37 | 37/37 | 37/37 | 37/37 |
| B shuffled tools | 49/51 = 96% | 35/37 | 34/35 | 34/34 | 34/34 | 34/34 |
| D inverted tool count | 51/51 = 100% | 37/37 | 36/37 | 36/36 | 36/36 | 36/36 |
| C base, no adapter | 27/51 = 53% | 17/37 | 17/17 | 17/17 | 17/17 | 17/17 |
Read the cascade left to right: each column is computed only over the cases that survived the
previous one, so the denominators shrink. emitted <tool_call> counts, out of the 37 questions
that required a tool, how many produced a call that could be read at all. It is the gate that
makes the columns to its right comparable. C scores right tool 17/17 not because the base
model always picks the right tool, but because it only got that far 17 times, on the cases it
found easy. A model that attempts only when confident scores well on a sample it selected
itself. The right tool, params ok and types ok checks only apply once a call exists, so
they say nothing about the calls that were never made.
The reading, in three levels
Same data, fixed denominators. This is the honest comparison.
| base | with the adapter | |
|---|---|---|
| produced a usable call when it should have (37 cases) | 17/37 = 46% | 37/37 = 100% |
| delivered a usable and correct answer (51 cases) | 27/51 = 53% | 51/51 = 100% |
| decided right among the times it formatted | 27/29 = 93% | 51/51 = 100% |
Dataset
Synthetic traces in Portuguese, generated via Groq by two models:
llama-3.3-70b-versatile: the 30 fictional tools, the 518 user questions, and the answers for thesem_toolcases.qwen/qwen3.6-27b: the tool calls, the simulated API returns, and the final answers of the traces that do use a tool.
annajuliaasf/tool-use-traces-ptbr
Reproduction
The whole pipeline (query generation, trace generation, preparation, training and evaluation) is at https://github.com/annaferreiras/ft-llm-tool-calling
Credits
- Trained with Unsloth, which is what made the project
viable on a free T4:
FastModelwithload_in_4bit=True,get_peft_modelfor the LoRA,train_on_responses_onlyso the loss counts only on the assistant's turns, andfinetune_vision_layers=False, which dissolves into a single parameter the complication of Gemma 4 being multimodal.
The other references used along the way are in the repository README.
- Downloads last month
- 23