YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

BFCLv4 Submission Repository

Complete submission repository for Berkeley Function Calling Leaderboard (BFCL) v4.

Repository Contents

  • handler.py - Main entrypoint with handle_request() function
  • requirements.txt - Python dependencies
  • MODEL_CARD.md - Model information and training details
  • inference_example.py - Simple usage example
  • train_lora.sh - LoRA fine-tuning script (GPU required)
  • tests/ - Test suite with fixtures
  • .github/workflows/ci.yml - CI pipeline
  • cloud_build_instructions.md - Cloud deployment guide
  • bfcl_submission_instructions.md - Submission checklist

Environment Variables

  • MODEL_NAME - HuggingFace model identifier (default: meta-llama/Llama-2-13b-chat)
  • BFCL_MODE - Operation mode: prompt or fc (default: prompt)
  • DEVICE - Compute device: cpu or cuda (default: cpu)
  • HF_TOKEN - HuggingFace token for gated models (optional)

Local Setup

Important: Use Python 3.10 or 3.11 (PyTorch doesn't support 3.13 yet)

python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Running Inference

Basic usage (uses mock responses for testing):

python inference_example.py

With a real model:

# TinyLlama (2.2GB download)
MODEL_NAME="TinyLlama/TinyLlama-1.1B-Chat-v1.0" DEVICE="cpu" python inference_example.py

# Larger model on GPU
MODEL_NAME="meta-llama/Llama-2-13b-chat" DEVICE="cuda" python inference_example.py

Running Tests

Quick test (uses mock responses):

pytest -q

With real model:

MODEL_NAME="TinyLlama/TinyLlama-1.1B-Chat-v1.0" DEVICE="cpu" pytest -v

Function Calling Modes

Prompt Mode (default)

Constructs a prompt with available functions and generates output containing FUNCTION_CALL: {"name": "...", "args": {...}}. The handler parses this structured output.

Native FC Mode

Set BFCL_MODE=fc to use native function calling APIs if supported by the model.

Response Format

{
  "status": "ok",
  "response": {
    "type": "function_call",
    "name": "function_name",
    "arguments": {...}
  },
  "raw_model_output": "..."
}

Parallel Function Calls

The handler processes requests that may require multiple function calls. The current implementation returns a single function call per request. Multi-call scenarios are handled through iterative turns.

Testing

The test suite covers:

  • Simple function calls with typed parameters (simple_python)
  • Parallel function execution patterns
  • Multi-turn conversations with missing parameters
  • Memory KV operations (memory_kv)
  • Web content fetching and extraction (web_search_no_snippet)

All tests use local fixtures and run deterministically. When no real model is loaded, the system uses intelligent mock responses that extract function information from the available functions list.

Training

LoRA fine-tuning requires GPU:

bash train_lora.sh

Do not run on CPU-only machines.

License

Apache-2.0

Downloads last month
5
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support