Developed by: cstr
License: apache-2.0
Finetuned from model : vonjack/Phi-3-mini-4k-instruct-LLaMAfied

This is a quick experiment with only 150 orpo steps from a german dataset.

This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.

Downloads last month: 4

Inference Examples

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Model tree for cstr/phi-3-orpo-v8_16

Base model

vonjack/Phi-3-mini-4k-instruct-LLaMAfied

Finetuned

(4)

this model

Finetunes

1 model

Quantizations

2 models