MiMo-9B-Cold (LIMA DPO)

A cold-style preference-aligned model built on MiMo-V2.6-Distill-Qwen-9B: DPO applied directly to the base model (no SFT), using 1,027 modified LIMA-style preference pairs to make responses concise, decisive, and direct — while preserving intelligence and general knowledge.

📋 Model Overview

Attribute Value
Model Name wesjos/Mimo-Qwen3.5-9B-Cold
Base MiMo-V2.6-Distill-Qwen-9B (Qwen3.5-architecture distill, 9.4B params)
Method DPO (Direct Preference Optimization), skipping SFT
Training Data lima_dpo_clean.jsonl — 1,027 LIMA-style preference pairs
Training Framework Unsloth + TRL DPOTrainer (QLoRA 4bit)
Release Formats HF merged_16bit (18 GB) · LoRA adapter · GGUF Q8_0 (8.87 GB)
Context Length 262,144 (native)
Languages English / Chinese

🚀 Usage

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "<repo>/merged_16bit", torch_dtype="bfloat16", device_map="auto"
)
tok = AutoTokenizer.from_pretrained("<repo>/merged_16bit")

messages = [{"role": "user", "content": "Introduce yourself in two sentences."}]
inputs = tok.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    enable_thinking=False, return_tensors="pt", return_dict=True,
).to("cuda")
out = model.generate(**inputs, max_new_tokens=300, temperature=0.6, top_p=0.9)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

llama.cpp (GGUF Q8_0)

llama-server -m mimo-9b-cold-Q8_0.gguf \
  --ctx-size 8192 -ngl 99 -ctk q8_0 -ctv q8_0 -fa on \
  --jinja --alias mimo-9b-cold --port 18080

Recommended Parameters

  • temperature: 0.6, top_p: 0.9, top_k: 20
  • Anti-repetition: repeat_penalty 1.05, DRY multiplier 0.8
  • Thinking mode: template supports enable_thinking toggle

⚠️ Limitations & Notes

  1. Small training scale: only 1 epoch × 1,027 pairs — lightweight style alignment; do not expect capability gains, only style transfer
  2. Slightly higher tool hallucination rate: for function calling, add parameter schema validation at the application layer
  3. Knowledge cutoff: inherited from the base model (Qwen3.5 family); training data contains no new knowledge
  4. Safety alignment: retains the base model's full safety values (illegal requests are refused with lawful alternatives)
  5. BBH sampling variance: at limit 200 each subset has only ~4 questions; the -2.78pp drop is not statistically significant

🙏 Acknowledgements

  • Base: MiMo-V2.6-Distill-Qwen-9B (Qwen3.5 architecture)
  • Methods: LIMA (Less Is More for Alignment) · DPO (Rafailov et al.)
  • Tools: Unsloth · TRL · evalscope · llama.cpp

Downloads last month
405
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wesjos/Mimo-Qwen3.5-9B-Cold

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(25)
this model