Orbit Support Drafter (DPO)

The Drafter agent of a LangGraph multi-agent support system for the fictional shop Orbit Electronics. Base model TinyLlama/TinyLlama-1.1B-Chat-v1.0, LoRA-fine-tuned and merged. Stage: dpo (sft = supervised on the response simulator; dpo = then aligned with human and critic preferences via Direct Preference Optimisation).

Evaluation (held-out simulated tickets, full multi-agent graph)

stage temperature first-draft approval resolved avg drafts
base 0 0 0 3
sft 0 1 1 1
sft 0.7 1 1 1
dpo 0 1 1 1
dpo 0.7 1 1 1

Knowledge-base fingerprint: 034b279610ce. Demo model trained on synthetic data; it only knows the toy policies of Orbit Electronics.

Preferences: {'synthetic': 24, 'critic': 14, 'gold': 2} (β=0.1, epochs=2).

Downloads last month
118
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for trychoosing/orbit-drafter-tinyllama-1.1b

Adapter
(1596)
this model