EXP-aicomps-mm20 โ SFT vs SDFT under 20% mismatched demonstrations
Research checkpoints: Qwen/Qwen2.5-7B-Instruct fully fine-tuned on the ToolUse task with SFT or SDFT, on clean demonstrations or with 20% of demonstrations deliberately mismatched to other questions (mm20). Only the seed-42 final checkpoint of each of the 4 configurations is published; each lives in its own subfolder with a model card, sha256 list and eval results.
| Subfolder | Method | Training data |
|---|---|---|
tooluse-sft-clean-seed42 |
SFT | clean |
tooluse-sft-mismatch-seed42 |
SFT | 20% mismatched |
tooluse-sdft-clean-seed42 |
SDFT | clean |
tooluse-sdft-mismatch-seed42 |
SDFT | 20% mismatched |
mismatch models were trained on intentionally corrupted data and are for research only.
PROTOCOL.md is the locked experiment protocol. manifests/ holds the train/val split (record ids) and the corruption mapping (record indices + hashes; no dataset text).
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support