YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
BFCLv4 Submission Repository
Complete submission repository for Berkeley Function Calling Leaderboard (BFCL) v4.
Repository Contents
handler.py- Main entrypoint withhandle_request()functionrequirements.txt- Python dependenciesMODEL_CARD.md- Model information and training detailsinference_example.py- Simple usage exampletrain_lora.sh- LoRA fine-tuning script (GPU required)tests/- Test suite with fixtures.github/workflows/ci.yml- CI pipelinecloud_build_instructions.md- Cloud deployment guidebfcl_submission_instructions.md- Submission checklist
Environment Variables
MODEL_NAME- HuggingFace model identifier (default: meta-llama/Llama-2-13b-chat)BFCL_MODE- Operation mode:promptorfc(default: prompt)DEVICE- Compute device:cpuorcuda(default: cpu)HF_TOKEN- HuggingFace token for gated models (optional)
Local Setup
Important: Use Python 3.10 or 3.11 (PyTorch doesn't support 3.13 yet)
python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
Running Inference
Basic usage (uses mock responses for testing):
python inference_example.py
With a real model:
# TinyLlama (2.2GB download)
MODEL_NAME="TinyLlama/TinyLlama-1.1B-Chat-v1.0" DEVICE="cpu" python inference_example.py
# Larger model on GPU
MODEL_NAME="meta-llama/Llama-2-13b-chat" DEVICE="cuda" python inference_example.py
Running Tests
Quick test (uses mock responses):
pytest -q
With real model:
MODEL_NAME="TinyLlama/TinyLlama-1.1B-Chat-v1.0" DEVICE="cpu" pytest -v
Function Calling Modes
Prompt Mode (default)
Constructs a prompt with available functions and generates output containing FUNCTION_CALL: {"name": "...", "args": {...}}. The handler parses this structured output.
Native FC Mode
Set BFCL_MODE=fc to use native function calling APIs if supported by the model.
Response Format
{
"status": "ok",
"response": {
"type": "function_call",
"name": "function_name",
"arguments": {...}
},
"raw_model_output": "..."
}
Parallel Function Calls
The handler processes requests that may require multiple function calls. The current implementation returns a single function call per request. Multi-call scenarios are handled through iterative turns.
Testing
The test suite covers:
- Simple function calls with typed parameters (simple_python)
- Parallel function execution patterns
- Multi-turn conversations with missing parameters
- Memory KV operations (memory_kv)
- Web content fetching and extraction (web_search_no_snippet)
All tests use local fixtures and run deterministically. When no real model is loaded, the system uses intelligent mock responses that extract function information from the available functions list.
Training
LoRA fine-tuning requires GPU:
bash train_lora.sh
Do not run on CPU-only machines.
License
Apache-2.0
- Downloads last month
- 5