BitAgent-Bounty-8B-dpo-prune

Fine-tuned from BitAgent/BitAgent-Bounty-8B (Apache-2.0) with iterative step-level DPO for multi-turn tool-calling efficiency.

Changes from the base model

  • Method: iterative step-level DPO + validated trajectory pruning for multi-turn tool-use efficiency (merged weights from LoRA fine-tuning).

  • Training signal: self-generated preference pairs on BFCL v3 Base Multi-Turn.

A detailed description of the method and experiments will appear in an upcoming paper.

License

Apache-2.0. This is a derivative of BitAgent-Bounty-8B and is not an official BitAgent release. Please retain the original attribution when redistributing.

Intended use

Multi-turn function calling / tool use evaluation (e.g., BFCL).

Downloads last month
245
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lwa201/BitAgent-Bounty-8B-dpo-prune

Finetuned
(1)
this model