Your paragraph text

Qwen3-4B-Instruct-2507 Distilled from Kimi K2

This is a repository which hosts the model files of the Model Distillation performed from the parent model Kimi K2 to the student model Qwen3-4B-Instruct-2507, with the aim of further advancing the student's agentic capabilities, particularly multi-step tool calling.

Model Overview

Parent (teacher) model Kimi K2
Student model Qwen3-4B-Instruct-2507
Dataset Agent-Ark/Toucan-1.5M (Kimi-K2 configuration)
Method Supervised fine-tuning on teacher trajectories

Training Data

Toucan-1.5M contains over 1.5 million agentic trajectories synthesized from 495 real-world Model Context Protocol (MCP) servers, spanning more than 2,000 tools. It includes single-turn and multi-turn interactions, as well as sequential and parallel tool calls with real tool execution.

Training used the Kimi-K2 configuration of the dataset. A total of 9,168 samples were used for a single epoch, taken from all four subsets in the following distribution:

Subset Samples Share
single-turn-original 2,750 30%
single-turn-diversify 2,292 25%
irrelevant 1,375 15%
multi-turn 2,751 30%

How to use it

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "Rumiii/Qwimi-4B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="cuda"
)

messages = [
    {"role": "user", "content": "What is the capital of France?"}
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=200,
    do_sample=False
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Only a portion of the full dataset was used. Samples longer than 4,096 tokens were excluded (no sample was truncated), as were multi-turn samples that contain tool calls but no declared tools. The trajectories follow the standard Qwen3 tool-calling chat template.

Note

This distillation was performed through an existing dataset. The Kimi K2 responses were not generated live for this model; the trajectories come from Agent-Ark/Toucan-1.5M, which was already available on Hugging Face.

Downloads last month
191
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rumiii/Qwimi-4B

Finetuned
(2360)
this model
Quantizations
2 models

Dataset used to train Rumiii/Qwimi-4B